Original Summary

Hey everyone, I'm currently building an AI-powered coaching and journaling app, and I'm stressing a bit about backend scalability and API costs. Right now, users interact with a custom-trained cognitive/MBTI framework, meaning there's a fair amount of context, memory state, and prompt engineering going on per session. If the app grows, relying solely on heavy commercial models via API could quickly eat up any subscription margins, especially at a $12-$15 price point. For those of you running production B2C AI apps, what is your go-to architecture for cost reduction ? Are you using smaller open-source models hosted locally or on cheaper cloud providers (like Groq, Together AI, etc.) for repetitive tasks? How do you handle caching or smart context pruning to avoid sending massive history tokens on every single message?   submitted by   /u/0x_akerue [link]   [comments]


  • 情报分类:商业与市场研究
  • 分类依据:内容涉及商业、投资或市场动态
  • 信息来源:Reddit · SaaS
  • 发布时间:2026/9/19 09:08:04