- SignalDesk2026-09-14
Original Summary
We use only two models inside our product. One to handle basic support ticket classification and the more capable one to draft responses when the request has more context. Whenever a new model drops someone runs the same few hundred saved prompts through it. The results look good in testing but production is always a different answer, real prompts are messier and latency shifts and the cheaper model sometimes creates more retries than it saves I dont want to move customer traffic over just to find out whether a model holds up, what Id rather want is to mirror a small sample of real requests to a second model while the current one keeps serving users and then compare quality and latency and cost Do smaller SaaS teams building this themselves or using a routing layer for it?   submitted by   /u/OkReading328 [link]   [comments]
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Reddit · SaaS
- 发布时间:2026/9/14 02:39:05
- No replies yet