- SignalDesk2 hr ago
Original Summary
The hardest part of autonomous coding agents wasn't getting them to code I've been building a system for running AI agents as a team, and one thing surprised me: the problem stopped being about prompting pretty quickly. A single coding agent can produce a decent patch. But that's not really how I want work to happen. Someone needs to understand the problem and plan it. Someone builds it. Someone else should review it. Sometimes security needs a separate look. And eventually something has to decide whether the work is actually ready to ship. The rule that ended up mattering most to me was: The agent that does the work shouldn't be the agent that decides whether the work is good enough. So I started treating agents less like chatbots and more like workers inside a workflow. A typical software workflow for me now looks something like: Plan → Implement → Code Review → Security Review → Approval → Merge Some steps can run in parallel. A reviewer can send the work back. A human can approve before merge, or the whole thing can run autonomously. That introduced a completely different set of problems. I needed to know what each agent did, what it decided, why something was rejected, which model was responsible for a step, what happened when an agent failed, and whether I could trust the system to merge something without me watching it. And once agents started working through longer workflows, another problem became obvious: context gets expensive very quickly. Simply passing the entire conversation and previous agent output to the next model works, but it wastes tokens and eventually becomes its own scaling problem. So I ended up treating context as infrastructure too. UstaWork doesn't just pass AI output from one agent to another. It uses things like context compression, memory and caching to reduce how much information needs to be sent again, and to avoid unnecessary model turns when the system already has what it needs. The goal isn't to make agents "forget" things. It's to keep the useful context while avoiding repeatedly paying for the same information. That matters quite a lot once a workflow has multiple agents reviewing and handing work back and forth. Better context management means fewer tokens, fewer unnecessary turns, and significantly lower model cost. So the interesting problem became less: "How do I make an AI write better code?" and more: "How do I govern a group of AIs doing work without babysitting them, while keeping the whole thing efficient enough to actually run?" That eventually became UstaWork . It's a control plane where you build workflows from specialist agents, reviews, conditions and approval gates. The agents do the work; the workflow determines who gets to judge it, what context they need, and whether the work can continue. Software engineering is the first use case, but I'm increasingly interested in whether this model makes sense for other delegated work too
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Reddit · SaaS
- 发布时间:2026/10/8 22:23:33
- No replies yet