Original Summary

Part of a longer side project (ClawHunt) put out a public model card for it. Fine-tuned on top of an abliterated Qwen 3.8-27B base, meant to make the model recall internal codebase conventions — migration rules, deploy checks, secret handling. r=16, ~80M trainable params, about 7.5 hours of training. Tried to keep the eval honest rather than make it look better than it is: Passed all 6 internal eval gates vs. the previous version, but the improvement wasn't statistically significant (McNemar p=0.5) Best score ties an earlier benchmark rather than beating it Ran it against Claude Sonnet 5 and Opus 5 on the same 96 prompts — our model wins on private-repo recall (expected, Claude never saw that data), but the two Claude models failed the hallucination-related gates in different ways from each other Still experimental, not deployed anywhere yet but Open Source of course. This is more of a "here's a useful pattern if you're doing something similar" share than a finished product. One thing worth knowing: the base model is an abliterated build, so refusal behavior is different from stock Qwen. Didn't add that ourselves, it just comes with the base, and it's noted in the card so anyone building on top of it knows what they're starting from. Model card for reference: huggingface.co/Clawhunt-store/clawhunt-p6-candidate-20260909   submitted by   /u/Similar_Job_6080 [link]   [comments]


  • 情报分类:商业与市场研究
  • 分类依据:内容涉及商业、投资或市场动态
  • 信息来源:Reddit · SideProject
  • 发布时间:2026/9/18 18:09:19