我操,太整活了! 小米 MiMo 团队居然在直播他们新模型 MiMo-V2.6 的RL 训练过程,看起来是一直在进行自动化的评估和测试。 然后一个特别的点是,它显示了成本,而且成本是实时显示的。 你可以看到训一个模型到底有多费钱,每一秒都在涨那个钱! Fuli Luo : Nearly half a year of silence. We spent it studying one problem: how far RL can scale. MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task Link: https://x.com/_LuoFuli/status/2100296686719610932


  • 情报分类:技术学习与提效
  • 分类依据:内容涉及技术、AI、软件工具或工程实践
  • 信息来源:AI / AI前沿 / 歸藏 @op7418
  • 发布时间:2026/9/17 10:44:44