Original Summary

PicoLM is an LLM inference engine written in C99. It currently supports llama-2, GPT-2, Qwen 3.6&#x2F;3.8(+MoE) and Gemma-3n models. Significant amount of work went into CPU SIMD acceleration&#x2F;testing&#x2F;correctness, and wide cross-platform availability with constant testing to never lose portability (from DOS through OS&#x2F;X 10.4 to modernity). CUDA&#x2F;HIP is supported, and accelerated IMMA kernels are available. More work needs to be done on prompt processing speed, but text generation is quite fast already.<p>With v1.0-rc2, the list of tested and supported platforms became even more obscene (iPhone 1! Tru64!), and some amount of work has gone into having a Vulkan backend.


  • 情报分类:硬件与数码
  • 分类依据:内容涉及硬件、数码产品或通信卡
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/9/24 00:12:19