- SignalDesk3小时前
Original Summary
Open weight AI is like open source software. Users not only run the weights, but also modify the weights. It matters to develop training framework for local hardware.<p>The method to train large AI models on local hardware is generally called QLoRA (Low Rank Adaptation over Quantized base model). In the last few years it's usually done with HuggingFace Transformers (which is the basis of training frameworks such as Unsloth and Axolotl) and bnb 4-bit base model. However, bnb does not yet support MoE models, so the local training of recent MoE models seemed stall for some time.<p>Since Transformers 5.18, it's started to support GGUF, and it supports recent models with MoE and sparse attentions (and it's not slow, already faster than llama.cpp on Mac). GGUF is a versatile container format. It can be smaller than 4 bpw with surprisingly good quantization quality, so there are new possibilities for local training.<p>I've shown that we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM without CPU offload. I've optimized it on Strix Halo and it trains at 200 token/s. There is still room to optimize, compared to > 1600 token/s prompt processing we've achieved, and the common sense that LoRA training (with gradient checkpointing) takes 4-5x work of prompt processing. CPU/disk offload (like Strata) and multi-GPU also need more work that I'm not currently focusing on.<p>On Strix Halo we can also train DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM at 100 token/s, but I think it's less practical than Qwen3.8FN for local use.
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Hacker News 新项目
- 发布时间:2026/10/11 00:43:43
- 暂无回复