Original Summary

Hey all, just wanted to share how my spare time has been spent for quite some time now.<p>Ok, this one started oddly. My aim initially was just to serve multiple models on one GPU and fire it up in my unraid machine.<p>By the time that first bit was done, I realised I could just add one more thing... and another... and well, from beginning of this year when, imho, AI became useful, the whole thing it kind of blew up.<p>So here&#x27;s what it does:<p>- spins a ray cluster<p>- deploys models using various loaders (vllm, llamacpp, diffusers, sherpa_onnx, etc.)<p>- serves an openai responses ready api on top of them<p>Now some the details, already listed in my production readiness (<a href="https:&#x2F;&#x2F;docs.model-ship.ai&#x2F;production-readiness&#x2F;#production-readiness" rel="nofollow">https:&#x2F;&#x2F;docs.model-ship.ai&#x2F;production-readiness&#x2F;#production-...</a>) plan:<p>- scores 17&#x2F;17 on the Open Responses conformance suite and it&#x27;s got its own responses redis backed store<p>- currently supports cpu, metal and cuda<p>- it has prometheus metrics and grafana dashboard<p>- can run under docker, k8s or simply native install<p>- automatically sizes model context<p>- has MCP support<p>Feedback would be much appreciated.<p><a href="https:&#x2F;&#x2F;github.com&#x2F;modelship-ai&#x2F;modelship" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;modelship-ai&#x2F;modelship</a>


  • 情报分类:商业与市场研究
  • 分类依据:内容涉及商业、投资或市场动态
  • 信息来源:Hacker News 新项目
  • 发布时间:2026/10/11 15:28:20