- SignalDesk2 hr ago
Original Summary
TL;DR: Augment Code swapped its production coding-agent backend in September. It didn't switch to a bigger model. It switched to a smaller one running on an architecture that doesn't generate tokens one at a time — Stefano Ermon's Mercury 2.5 — and the numbers aren't close: latency down 82%, cost down 90%, in a shipped product, not a benchmark deck. Diffusion produces a block of tokens in parallel. That's the exact bottleneck you've been fighting at 18% GPU utilization at 11:47pm, while yesterday's KV-cache tuning bought you another four percent. This isn't a paper anymore — Artificial Analysis independently measured it at 770 tokens/second against Inception's own claim of 1,107, and even the conservative number is very probably real, deployed, load-bearing. What it doesn't do is make you safer. Every governance document sitting on a desk somewhere — the RSP, the Preparedness Framework, the EU AI Act's Annex III, a workforce bill that stalled without a floor vote — regulates what a model may output. None of them regulate what happens to the engineer whose entire expertise sits on the side of the stack that just got walked past. THE GAP: Nobody's built the neutral test — a place to run your own traffic against Mercury-style diffusion and your current autoregressive stack, on the same box, with nobody grading its own homework. Artificial Analysis's Optima and SemiAnalysis's InferenceX get close, but neither runs both architectures side-by-side on your own data at your own concurrency. Open. Viable — a $417/month Optima seat and a $40M raise for a benchmarking startup called Vals both say the money's already moving nearby, and the resource bar is realistic: DiffusionGemma ships open-weight, on-demand H100s run $3.41/GPU-hour. The fastest way in isn't beating an incumbent — it's being first with an open, published, same-backbone result. Gemma 4 26B-A4B against its diffusion sibling, DiffusionGemma, same GPU, same vLLM, method public. That's the exact slot InferenceX, Artificial Analysis, and LMArena each took by being first and open — and the sellers end up citing the referee back (Inception's own Mercury paper cites Artificial Analysis). Profitability Horizon: not estimable from available research on its own — but on the Specification/Roadmap laid out under FEASIBILITY below, the first paid verdict clears its own compute cost inside that first job (roughly $34–$68 in GPU time against a $417/month seat as the price anchor). _________________________________________ FEASIBILITY: Opportunity: Open gap — no product meets all six requirements for an independent, own-traffic autoregressive-vs-diffusion test on your own box. Viable: adjacent products are already priced and selling. Specification: Tracks your own traffic sample, both architectures on identical hardware, latency and GPU utilization at your own concurrency, cost per completed task, quality on your own rubric, method p
- 情报分类:工作与职业机会
- 分类依据:内容涉及招聘、求职或职业发展
- 信息来源:Reddit · SideProject
- 发布时间:2026/9/25 11:30:10
- No replies yet