- SignalDesk1小时前
Original Summary
This model makes calibrated, typed decisions about an image plus optional text. It answers choice, score and noul (yes/no probability) questions in one forward pass, with no text generation.<p>It adds image input to Laya by replacing Laya's ModernBERT encoder with SmolVLM-256M-Instruct. Laya's predict(state, questions) API, proper-scoring-rule training and temperature calibration are unchanged.<p>check out the live demo at <a href="https://huggingface.co/spaces/thaitea/laya-vision-demo" rel="nofollow">https://huggingface.co/spaces/thaitea/laya-vision-demo</a><p>and the source<p>- <a href="https://huggingface.co/thaitea/laya-vision-smolvlm-256m" rel="nofollow">https://huggingface.co/thaitea/laya-vision-smolvlm-256m</a> - <a href="https://github.com/r33drichards/laya-vision" rel="nofollow">https://github.com/r33drichards/laya-vision</a>
- 情报分类:技术学习与提效
- 分类依据:内容涉及技术、AI、软件工具或工程实践
- 信息来源:Hacker News 新项目
- 发布时间:2026/9/19 23:37:34
- 暂无回复