Original Summary

Sorry if this is a stupid question but...<p>Problem:<p>I boot opencode into a codebase. I ask it to do one thing, it has to call 40 tools to &quot;remember&quot; what the codebase looks like relevant to that.<p>It has no memory so every fresh boot is re-explore.<p>My understanding, as an ML n00b, is the GPU have KV states of massive matrices that are multi GB in RAM (e.g., 10-30GB) so dumping to disk and pushing over network (<i>his shoulders shuddered at the thought of egress fees</i>), is not great.<p>So instead we have cache during &quot;work time&quot;, but off &quot;work time&quot;, we have to reboot and build the cache again.<p>But couldn&#x27;t CF or someone setup CDNs to cache token state between days or something like that?<p>Ages ago I was trying to solve part of this problem client side w&#x2F; vectoring a codebase and creating a tool to do quick queries, but I kind of stopped.<p>Curious if LLM token CDNs are gonna be a thing or not


  • 情报分类:商业与市场研究
  • 分类依据:内容涉及商业、投资或市场动态
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/9/16 00:20:29