Original Summary

Do you think we&#x27;re trading off model size, and model capacity, for compute? Take looped transformers, which GPT-6 Astra is rumored to be based on. They increase compute while keeping the model&#x27;s capacity almost fixed.<p>So my question is, if we had an infinitely large model (with also infinite capacity) and of course infinite compute, would we still need CoT or looped transformers? Could we have a model where the reasoning happens directly within a single forward pass? I think so. What we do today is just a trade-off imposed by the computing power we have available.


  • 情报分类:技术学习与提效
  • 分类依据:内容涉及技术、AI、软件工具或工程实践
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/10/8 17:15:25