- SignalDesk22小时前
Original Summary
I’ve been working on EvalSeal, a small open-source tool for making LLM eval results more trustworthy. The idea is simple: a single eval score is not enough. EvalSeal runs each case multiple times, measures flip rates, captures provenance, and seals the result into a tamper-evident ledger. A finding from the current repo: Using an LLM judge, 5 of 20 borderline cases flipped verdicts across repeated runs. Using numeric answer matching on 40 GSM8K cases, 0 cases flipped. Same model family, different grading method. The instability came from the judge, not the target model. Current release is v0.3.0. I’m working on the v1 path now with signing, suite files, and benchmark-oriented scoring. Would love feedback from anyone working on LLM evals, CI gates, or agent reliability.   submitted by   /u/Fit_Fortune953 [link]   [comments]
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Reddit · SideProject
- 发布时间:2026/9/18 16:08:15
- 暂无回复