Original Summary

Hi guys,<p>A few months ago, I&#x27;ve been training several simple models as a side hobby. I tried training N models at once on a single GPU, but OOM spikes kept crashing my runs. So I built my own simple way to manage this kind of training. At the time, I also wanted to make some sort of interface that AI agents could use to dynamically allocate GPU resources and train models on their own. It&#x27;s been quite fun working on this, so I wanted to share it in case any of you are looking for something similar. Claude helped me write most of the code, but I&#x27;ve made sure that the stuff works as described (I&#x27;ve used this exact repo for my own training runs).<p>Hope you like it!


  • 情报分类:工作与职业机会
  • 分类依据:内容涉及招聘、求职或职业发展
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/9/27 16:08:41