Original Summary

Open weight AI is like open source software. Users not only run the weights, but also modify the weights. It matters to develop training framework for local hardware.<p>The method to train large AI models on local hardware is generally called QLoRA (Low Rank Adaptation over Quantized base model). In the last few years it&#x27;s usually done with HuggingFace Transformers (which is the basis of training frameworks such as Unsloth and Axolotl) and bnb 4-bit base model. However, bnb does not yet support MoE models, so the local training of recent MoE models seemed stall for some time.<p>Since Transformers 5.18, it&#x27;s started to support GGUF, and it supports recent models with MoE and sparse attentions (and it&#x27;s not slow, already faster than llama.cpp on Mac). GGUF is a versatile container format. It can be smaller than 4 bpw with surprisingly good quantization quality, so there are new possibilities for local training.<p>I&#x27;ve shown that we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM without CPU offload. I&#x27;ve optimized it on Strix Halo and it trains at 200 token&#x2F;s. There is still room to optimize, compared to &gt; 1600 token&#x2F;s prompt processing we&#x27;ve achieved, and the common sense that LoRA training (with gradient checkpointing) takes 4-5x work of prompt processing. CPU&#x2F;disk offload (like Strata) and multi-GPU also need more work that I&#x27;m not currently focusing on.<p>On Strix Halo we can also train DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM at 100 token&#x2F;s, but I think it&#x27;s less practical than Qwen3.8FN for local use.


  • 情报分类:商业与市场研究
  • 分类依据:内容涉及商业、投资或市场动态
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/10/11 00:43:43