- SignalDesk2 hr ago
Original Summary
Hey all, just wanted to share how my spare time has been spent for quite some time now.<p>Ok, this one started oddly. My aim initially was just to serve multiple models on one GPU and fire it up in my unraid machine.<p>By the time that first bit was done, I realised I could just add one more thing... and another... and well, from beginning of this year when, imho, AI became useful, the whole thing it kind of blew up.<p>So here's what it does:<p>- spins a ray cluster<p>- deploys models using various loaders (vllm, llamacpp, diffusers, sherpa_onnx, etc.)<p>- serves an openai responses ready api on top of them<p>Now some the details, already listed in my production readiness (<a href="https://docs.model-ship.ai/production-readiness/#production-readiness" rel="nofollow">https://docs.model-ship.ai/production-readiness/#production-...</a>) plan:<p>- scores 17/17 on the Open Responses conformance suite and it's got its own responses redis backed store<p>- currently supports cpu, metal and cuda<p>- it has prometheus metrics and grafana dashboard<p>- can run under docker, k8s or simply native install<p>- automatically sizes model context<p>- has MCP support<p>Feedback would be much appreciated.<p><a href="https://github.com/modelship-ai/modelship" rel="nofollow">https://github.com/modelship-ai/modelship</a>
- 情报分类:商业与市场研究
- 分类依据:内容涉及商业、投资或市场动态
- 信息来源:Hacker News 新项目
- 发布时间:2026/10/11 15:28:20
- No replies yet