- SignalDesk1小时前
Original Summary
I'm tired of AI coding agents which are good for 30 seconds and then completely fail on the slightest hiccup: a failing test, a temporarily unavailable dependency, an incorrect assumption. Sometimes I end up having to babysit these agents anyway.<p>Relay is a harness for coding agents which takes advantage of the fact that agents are often good at doing something slightly wrong, and not so good at doing something correctly and completely. Relay repeatedly tries the task and on each failure, uses the output to find a better way to do it.<p>It does this by running the task in a loop: attempt the task, run the checks in a fresh sandbox, on failure read the error and retry instead, until the checks pass. What it learned from a run is retained, so that it doesn't repeat the same mistake, and when the checks eventually pass, it will ship the change as a normal git PR - no hidden state, no magic, the diff is human readable.<p>A few specifics:<p>- it runs locally, is free, and doesn't require an account to run the local agent - it uses a top-tier model to plan/review, and cheaper models to actually write code (for cost reasons) - happy to discuss why this is a good idea and why it isn't - the checks run in an isolated sandbox for each task, rather than "trusting" the agent - currently supports [list actual models/integrations - Opus 5.5, GPT-6, etc] - there's a paid tier for running this unattended across multiple repos with shared sandbox minutes, the actual agent is free.<p>It's not magic, it's not AGI, and it will fail badly on some problems.<p>Would appreciate feedback, particularly from people who've deployed agents unattended and found that they don't actually work, or have had to build substantial tooling around them to get useful results.
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Hacker News 新项目
- 发布时间:2026/9/30 02:46:07
- 暂无回复