Original Summary

I'm building a security product called Velmère and recently ran into a problem I think a lot of technical SaaS founders probably face. Our internal tests kept getting better. More regressions passed, more edge cases were covered and the dashboard looked increasingly green. That sounds good until you ask a different question. Are we actually improving the product, or are we just getting better at passing tests we wrote ourselves? So I changed how we're evaluating it. Instead of using only internal cases, we started testing against pinned public smart contract datasets. The first SmartBugs checkpoint gave us 29 relevant cases in vulnerability families supported by the current engine. 22 could be executed properly with the required compiler and AST evidence. All 22 were detected. The other 7 weren't quietly counted as safe. The system withheld the result because it couldn't prove enough to make the call. We also tested 29 OpenZeppelin 5.0.2 control candidates and 2 produced alerts. I'm not calling those false positives until someone independently validates the controls. And I'm not calling 22/22 “100% accuracy” either. The next step is deliberately worse for the marketing numbers: run the full corpus, include unsupported categories in the denominator and then replay historical code that was already independently audited. I'd rather have a benchmark that exposes where the product is weak than one that gives me a pretty screenshot.   submitted by   /u/Velmere_ [link]   [comments]


  • 情报分类:商业与市场研究
  • 分类依据:内容涉及商业、投资或市场动态
  • 信息来源:Reddit · SaaS
  • 发布时间:2026/9/21 07:43:37