Original Summary

General purpose LLMs, and classifiers like Jev, are nice and easy to use right away. That makes the pain of labelling data and creating specialized models for simple tasks not seem worth it, but for a project I am working on, I needed to create a large number of specialized models that could run efficiently on-device.<p>This project started as a way for me to test if GEPA could get frontier LLMs to generate better labels for me. I had mixed results. It works better on weaker models than on true frontier models, but there were definitely some gains.<p>I decided I&#x27;d rerun some of the flows on public datasets that are often used for these comparisons and post it, as I&#x27;m curious 1) what results others get with it, and 2) what ideas others have for doing a better job of this.<p>I&#x27;ve tried to document everything I could thoroughly, but always happy to chat.<p>Some things I want to try next: 1) Generating synthetic questions to train on, not just the labels, likely by using multiple models to validate agreement on whether it&#x27;s a worthwhile question to add to the set 2) Support RAG in the labelling flow by simply pointing to some docs or a corpus and have the rest be automated<p>I couldn&#x27;t think of a good project name so shrewd is just: shrew-&gt;something small, and d-&gt; distillation. I know distillation is a bit of a loaded word right now, but this is more just a labelling task, and into very small single-purpose models, so it is hopefully not an issue.


  • 情报分类:工作与职业机会
  • 分类依据:内容涉及招聘、求职或职业发展
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/9/23 04:00:46