Original Summary

New benchmark from researchers at Meta, Stanford, Harvard, UW, including the researchers who&#x27;ve worked on SWE-bench, ProgramBench etc.<p>Most benchmarks just test if AI can fix a problem you&#x27;ve already pointed out.<p>But obviously it would be much better to fix problems before you or any user runs into it. Like, isn&#x27;t it crazy that we still have to wait for people to open tickets before a lot of obvious bugs get found?<p>We wanted to test that capability at scale. Turns out that models still are terrible at it (best setup we tested still fixed &lt; 5% of bugs)<p>We have 100 repos of 22 languages and 4k bugs between. All the bugs are real-world bugs from github. We do a lot of filtering to ensure everything can be solved in this setting.<p>Model Resolve Cost Sol 5.6 (xhigh) 4.7% $7,230 Luna 5.6 (xhigh) 2.5% $224 Terra 5.6 (xhigh) 1.5% $357 Luna 5.6 (high) 1.4% $28 Opus 5 (xhigh) 1.3% $5,363 Kimi K3 0.6% $2,451 Luna 5.6 0.5% $4 GPT-5.4 Mini (high) 0.5% $122 GPT-5.4 Mini 0.2% $5 Gemini 3.5 Flash Lite 0.1% $6<p>Also the best model is very expensive.<p>We have a lot more FAQ on the website <a href="https:&#x2F;&#x2F;swesweep.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;swesweep.com&#x2F;</a> Oh and we&#x27;re all open-source (MIT license) at <a href="https:&#x2F;&#x2F;github.com&#x2F;facebookresearch&#x2F;swe-sweep" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;facebookresearch&#x2F;swe-sweep</a><p>Curious what you all think!


  • 情报分类:技术学习与提效
  • 分类依据:内容涉及技术、AI、软件工具或工程实践
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/10/1 23:36:14