Original Summary

Hey HN! Yolanda and Spencer here - wanted to share a token compression tool that we’ve built for ourselves to save 30% costs on codex!<p>After maxing out sub and burning $700&#x2F;day per person on api, we fine tuned a compression model to trim codex&#x27;s tool call output to reduce input token + cache. It cut down tokens by 29.6% and now I just leave it on by default in Codex.<p>To avoid messing up w&#x2F; cache, we use proxy + fine tuned qwen model trained on preserving agent trajectory to remove tool call results before they go back to the model, leaving kv cache untouched.<p>The cli is free for everyone to use (<a href="https:&#x2F;&#x2F;github.com&#x2F;spenmcke&#x2F;compress" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;spenmcke&#x2F;compress</a>). Just lmk ur feedback and hacks to shave even more costs on astra! If you want to integrate it into your product to offer the best models at low cost, I can set you up with an sdk and api keys<p>PS: It’s built for coding agents, not conversational agents. I optimized it for file retrieval accuracy, trajectory preservation, and quality to get up to 30% cost reduction depending on how context-heavy the task is.<p>On security and privacy side, it&#x27;s a proxy wrapping your local codex and ZDR so it doesn&#x27;t retain any queries. It’s on by default in codex and when you don’t want compression, you can use codex --uncompress to disable it.<p>Give it a try: code is in <a href="https:&#x2F;&#x2F;github.com&#x2F;spenmcke&#x2F;compress" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;spenmcke&#x2F;compress</a><p>You can install the cli using<p>curl -fsSL <a href="https:&#x2F;&#x2F;install.everestagi.com&#x2F;install.sh">https:&#x2F;&#x2F;install.everestagi.com&#x2F;install.sh</a> | sh &amp;&amp; source ~&#x2F;.config&#x2F;everest&#x2F;shell.sh<p>Love to hear any feedback and learn your hacky ways to save token costs too!


  • 情报分类:技术学习与提效
  • 分类依据:内容涉及技术、AI、软件工具或工程实践
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/10/1 01:30:54