- SignalDesk1小时前
Original Summary
I trained a vision-language model that answers typed questions about an image. It uses also choice, score and noul, like Jev.<p>I thought, why Jev only processes text? There should be the possibility to process an image too. I will experiment with some pictures and questions in the next days to see, how well it performs in real life.<p>On my M1 Pro a request with six questionut 400 ms p95, about 60 ms on a desktop GPU. The server encodes the image once and scores each option as a short suffix against the KV<p>Looking forward to answer your questions! :)
- 情报分类:综合情报
- 分类依据:内容未命中明确的垂直分类规则,归入综合情报
- 信息来源:Hacker News 新项目
- 发布时间:2026/9/27 17:50:05
- 暂无回复