- SignalDesk2 hr ago
Original Summary
"Decision models" are a new kind of AI model that only does one thing: pick an answer from a fixed list, fast. Jev, Cloudflare's Clef and OpenAI's new Decisions API are the main ones. Every vendor publishes its own numbers, so I built a site that measures them on the same task and shows all the evidence. What it does: - Runs each model on the same 100 support messages ("which team should handle this ticket?") and reports accuracy, speed and estimated cost, with error margins. - Lets you open any of the 100 cases and see what each model answered. - Publishes the whole dataset and every answer for download (CC BY 4.0). What I found so far: the purpose-built models answer 5 to 12 times faster than a general LLM doing the same job, and the general LLM gets a few more tickets right. The fast models also fail in opposite ways: some route a vague message to a team when they should ask a question, and one asks too often. It's free and there is nothing to buy. It's my own project. https://decisionmodelhub.com/benchmarks/support-routing I'd like feedback on two things: is the results page understandable if you're not an ML person, and what would make you trust (or distrust) a benchmark like this?   submitted by   /u/AncientComment2352 [link]   [comments]
- 情报分类:工作与职业机会
- 分类依据:内容涉及招聘、求职或职业发展
- 信息来源:Reddit · SideProject
- 发布时间:2026/10/9 20:41:48
- No replies yet