Original Summary

Claude Code usage goes fast, and a lot of mine was going into work that didn't need a frontier model: docstrings, mechanical refactors, filling in a module from a contract I'd already specified. I thought of this project as leveraging what free resources in AI we have and having opus be the mind that controls it all. I get the same functionality with the dramatic decrease in cost. This runs those tasks on a free or cheap OpenAI-compatible endpoint, then hands you the git diff to review. The worker can read, search and edit files in the repo; it cannot run shell commands or git, by design. The interesting part is that delegation is not obviously a win. You still pay for the spec and the review. So the benchmark measures the ratio: tokens to review a worker's diff versus tokens to just do the task inline. Median across 9 tasks: 85.6% savings, 9/9 passing. Caveats are in bench/README.md, including that the same model ran both arms, so the ratio isn't portable to another model without re-measuring. Stdlib-only Python, MIT. https://github.com/aayushpokhrel1/delegation-pipeline   submitted by   /u/therocl [link]   [comments]


  • 情报分类:技术学习与提效
  • 分类依据:内容涉及技术、AI、软件工具或工程实践
  • 信息来源:Reddit · SideProject
  • 发布时间:2026/10/3 05:44:18