- SignalDesk4小时前
Original Summary
Hey HN,<p>Today, we're launching selfbench.dev, an open-source tool that lets you create and run evals automatically from your PRs.<p>Every benchmark with sufficient trust eventually gets benchmaxxed (Goodhart's law) - the labs are incentivized to maximize their scores on that benchmark, which isn't predictive on whether it'll actually work within your setup. This has been a time-consuming process that only the largest companies can afford to do, so we built self-bench to fix this.<p>Self-bench uses agents to author / review Harbor environments generated from your PRs. You then approve every eval that the agent created, and can model/harness evals concurrently on sandboxes.<p>To save you money, we also allow you to connect your OpenAI and Claude subscriptions so you don't have to pay raw token costs.<p>We ran evals on several large codebases like<p>- Posthog (<a href="https://selfbench.dev/PostHog/posthog" rel="nofollow">https://selfbench.dev/PostHog/posthog</a>)<p>- Next.js (<a href="https://selfbench.dev/vercel/next.js" rel="nofollow">https://selfbench.dev/vercel/next.js</a>)<p>- Sentry (<a href="https://selfbench.dev/getsentry/sentry" rel="nofollow">https://selfbench.dev/getsentry/sentry</a>)<p>- Pi (<a href="https://selfbench.dev/earendil-works/pi" rel="nofollow">https://selfbench.dev/earendil-works/pi</a>)<p>- as well as other fantastic open-source projects (all on selfbench.dev)<p>From our evals:<p>- Kimi K3 and GLM 5.3 are almost always more expensive, yet less performant than models like GPT-6.1 Sol and Claude Opus 5.5, due to token efficiency.<p>- GPT-6 Luna is almost always the most effective "cost-efficient" model we've tested, not open-source models.<p>We'd love for you to try this on your codebase and let us know what you think!
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Hacker News 新项目
- 发布时间:2026/10/6 00:58:47
- 暂无回复