Original Summary

I&#x27;m not talking only about Fable level models, since we are obviously already close with stuff like Qwen.<p>But I&#x27;m also wondering about being able to run them on consumer-end hardware.<p>I remember using a local model 2-3 years ago and had to wait around 2-3 minutes for a basic answer to be printed. Now I&#x27;m running a &quot;thinking&quot; Qwen on a 16GB GPU and I&#x27;m able to do anything I&#x27;d do with Opus a couple months ago, at nearly the same speed. But that does use my entire VRAM and most of the RAM I have. No way I can also run a game or something else on the side.<p>But like how we went from bulky PCs to smartphones 1000x faster, and at the rate local models already improved, do you think we&#x27;ll ever be able to have the same kind of models running locally, on affordable hardware, on our phones or maybe our fridges?<p>Not saying we should use them on anything, that will be a question for later, but strictly thinking about capabilities.


  • 情报分类:商业与市场研究
  • 分类依据:内容涉及商业、投资或市场动态
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/9/23 18:52:33