Original Summary

I&#x27;ve had some strange results with Jev, and I question the &quot;probabilistic&quot; aspects of it. So I put OpenAI Decisions API against classic statistics experiments to see how it holds up:<p>(1) loaded coin with probability of heads biased towards 70%<p>(2) marble selection from a jar, with replacement; 5 red, 3 blue, 2 white<p>I tested both predicate and choice questions. I ran thousand trials against each experiment, and I also did an experiment where I change the order of choices, to see if it matters.<p>Summary:<p> Using predicate questions gave nearly perfect&#x2F;expected probability outcomes. E.g. for the coin toss, it sampled heads 70% of the time, and for the marble experiment, it sampled the red marble 50% of the time<p> Asking it to &quot;choose an outcome&quot; behaved differently from drawing randomly - if the true probability of a red marble draw was 50%, using Decisions API produced 86%, i.e. it picked the right marble but gave a significantly more biased weight on its choice<p>* Changing the choice order changes the probabilities! Moving the red marble from first to last choice changed its probability estimate from 86% to 73%<p>I have full summary of results here: https:&#x2F;&#x2F;gist.github.com&#x2F;acatovic&#x2F;6b31f0061603b3de97731a8a29576dbd<p>My question to you is: how do you trust and implement a &quot;System One&quot; style classifier like Decisions API, in your work?


  • 情报分类:技术学习与提效
  • 分类依据:内容涉及技术、AI、软件工具或工程实践
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/10/7 21:01:53