- SignalDesk2小时前
Original Summary
I use Claude Code's /goal a lot: you say what "done" means and the agent keeps working until a small judge model says the goal is met. Then I noticed the judge only reads the chat and never runs anything. I measured it on 4 tasks with hidden graders: it said "met" in 17 of 17 plain runs, and 8 of those were broken. So I built goalpost as a side project. It's a hooks plugin, you keep typing /goal like before. It freezes the goal into checkable criteria, blocks the stop until each one has a fresh passing check, protects existing tests from quiet edits, and runs a fresh-eyes auditor agent at the end. The video is two real runs on the same task, side by side: 28/34 hidden tests without it, 34/34 with it. It's not magic: it roughly doubles time and cost, and it didn't make the small model reliable. I'd love feedback on that tradeoff, and on whether the README explains it well enough. Free and MIT: https://github.com/syntaxixr/goalpost   submitted by   /u/Sanechka_SS [link]   [comments]
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Reddit · SideProject
- 发布时间:2026/10/7 15:55:50
- 暂无回复