- SignalDesk1 hr ago
Original Summary
We’ve been building Auker 1.0 at Reasonary AI. Instead of routing an entire task through one fixed model or workflow, Auker breaks it into subproblems and learns how to allocate models, agents, tools, and reasoning across them. After 116 optimization experiments, the best policy reached 80.35% on our MMMU-Pro evaluation at a total cost of $9.73—about 6× lower than its initial baseline cost at essentially unchanged accuracy. The interesting part for us was that the improvement did not come from simply using a more powerful model. It came from changing how intelligence was allocated. We published the methodology, experiment history, cost assumptions, and Pareto-frontier results here: https://reasonary.ai/blog/auker-rewrote-itself-and-surpassed-gpt-56-luna-to-redefine-a-benchmark-pareto-frontier-9de7d3 I’d appreciate feedback on two things: Is the benchmark explanation understandable? What real workload would you use to stress-test this approach? Try our product here: https://app.reasonary.ai/   submitted by   /u/SnooDingos8551 [link]   [comments]
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Reddit · SideProject
- 发布时间:2026/10/1 01:01:20
- No replies yet