Original Summary

It is a bit strange how new model is released and hour after there is commentators declaring it complete trash and embarrassment to the AI industry, or the best thing since sliced bread. Surely they have not had the opportunity to test the ins and outs of the model yet? Or do people just blindly trust benchmarks as if they were not pretty easy to manipulate, as research has shown quite a few times now? Or is it just all vibe?<p>So how do you measure how one model is better than another?


  • 情报分类:技术学习与提效
  • 分类依据:内容涉及技术、AI、软件工具或工程实践
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/10/7 00:07:27