- SignalDesk1小时前
Original Summary
I made PagedServe because I wanted to understand LLM serving internals properly instead of treating vLLM as magic It is an open source Python runtime with continuous batching prefix caching blockwise paged attention and an experimental Apple Metal kernel The blockwise path cuts temporary attention memory from 12 MB to 96 KB at 2048 tokens but the Metal path is still slower so there is plenty to improve Would genuinely love contributors who can help with benchmarks Metal profiling model compatibility docs or just finding where the architecture is dumb I will review every real issue and PR https://github.com/vermasarthak/pagedserve If you find it useful star it so more systems people can find it   submitted by   /u/Accomplished_Row1433 [link]   [comments]
- 情报分类:技术学习与提效
- 分类依据:内容涉及技术、AI、软件工具或工程实践
- 信息来源:Reddit · SideProject
- 发布时间:2026/9/24 04:13:25
- 暂无回复