- SignalDesk1小时前
Original Summary
I own ComputeAtlas.ai. It helps plan hardware for running AI locally, and I built it with AI-assisted coding. The GPU checker separates the model file, KV cache and runtime allowance. You choose a model, GGUF format, context length and number of concurrent conversations, then see the memory estimate and where it stops fitting. For example, Qwen3 8B at Q4_K_M with one 8,192-token conversation comes to about 6.81 GiB using the default 1 GiB runtime allowance. On an 8 GiB GPU with 10% reserved, the budget is 7.20 GiB. Longer context or more conversations can put it over that budget. It covers 13 model profiles across Qwen, Llama, Mistral and DeepSeek distills. It can also show GPUs that meet the estimate and open a starting parts list in Builder to check compatibility and power. The numbers are estimates, not GPU benchmarks. The detailed checker currently assumes one GPU and FP16 KV cache; it doesn't cover CPU offload, multiple GPUs or every model architecture. Sources and limits are on the page. No account is needed. Is it clear which settings change the result? If you've measured a different allocation, I'd like to see the model file, runtime version, context, cache type and GPU so I can investigate the difference. Try the 8 GiB example   submitted by   /u/Haunting_Honeydew123 [link]   [comments]
- 情报分类:技术学习与提效
- 分类依据:内容涉及技术、AI、软件工具或工程实践
- 信息来源:Reddit · SideProject
- 发布时间:2026/10/9 22:30:38
- 暂无回复