Original Summary

This model makes calibrated, typed decisions about an image plus optional text. It answers choice, score and noul (yes&#x2F;no probability) questions in one forward pass, with no text generation.<p>It adds image input to Laya by replacing Laya&#x27;s ModernBERT encoder with SmolVLM-256M-Instruct. Laya&#x27;s predict(state, questions) API, proper-scoring-rule training and temperature calibration are unchanged.<p>check out the live demo at <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;thaitea&#x2F;laya-vision-demo" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;spaces&#x2F;thaitea&#x2F;laya-vision-demo</a><p>and the source<p>- <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;thaitea&#x2F;laya-vision-smolvlm-256m" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;thaitea&#x2F;laya-vision-smolvlm-256m</a> - <a href="https:&#x2F;&#x2F;github.com&#x2F;r33drichards&#x2F;laya-vision" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;r33drichards&#x2F;laya-vision</a>


  • 情报分类:技术学习与提效
  • 分类依据:内容涉及技术、AI、软件工具或工程实践
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/9/19 23:37:34