- SignalDesk1小时前
Original Summary
So for a while now I've been building a thing I call Orchestra. I type a task, and it splits the work between a bunch of Claude and Codex agents, each on its own git branch, and at the end I get something I can merge or throw away. Every task first goes to Jev. It just answers questions. How big is this, what kind of work is it, does it need a plan, is it risky. Then the Arena picks a model for each step. Every model (Haiku, Sonnet, Opus, and GPT-6 Luna, Sol, Astra) has a win/loss record per type of work. It rolls weighted dice on those records and takes cost into account, so easy stuff goes to cheap models and hard stuff to the expensive ones. Every pick comes with a reason line The records come from what actually happens. Tests passing, reviews sending work back, who wins a duel, my thumbs up or down, and whether I merge the branch or quietly delete it. The judge in a duel only sees "A" and "B", never the model names. If the judge isn't sure, nobody learns anything from that round. It all goes into a SQLite file on my box, so it slowly learns what works on my own repos. I asked AI to fix AI slop; I typed "We need to build a new UI for Orchestra. Maybe even go back in time because everything is AI slop. It spun up 15 agents. Two design agents competed and the judge said neither of them did the job, confidence 0.19. It called in a stronger model. The plan said I had to approve the design first. I was busy recording this video, so the agent just restyled everything anyway. The reviewer caught it and wrote "auto mode does not satisfy the plan's gate." Snitched on a coworker. In the end Sonnet got the fix, cleaned up the stylesheet and the tests passed. 15 agents, 227 tool calls, 9 changed files, to redo some CSS. Totally worth it.   submitted by   /u/MidgetTower [link]   [comments]
- 情报分类:工作与职业机会
- 分类依据:内容涉及招聘、求职或职业发展
- 信息来源:Reddit · SideProject
- 发布时间:2026/9/30 05:50:21
- 暂无回复