- SignalDesk1小时前
Original Summary
the per request average was looking fine. the routing policy narrows 33 tools down to 4 candidates but all 33 schemas still land in the context window on every turn. I figured prefix caching was absorbing enough of that payload to make tool schema pruning something I could deal with later. not ideal. so I priced a typical 18 turn support session end to end and the schemas alone push roughly 72000 input tokens. the generation bill is lower than the schema block (I realized this while reviewing a single invoice line). I'm evaluating Braintrust for span level token attribution, experiment diffs across schema sets and regression datasets around routing changes so I can test pruning without breaking tool selection. has anyone found a good way to measure whether an experiment diff on smaller schema sets is preserving routing quality while cutting session cost?   submitted by   /u/Primary-Life-4291 [link]   [comments]
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Reddit · SaaS
- 发布时间:2026/9/29 03:39:35
- 暂无回复