Ok, this one started oddly. My aim initially was just to serve multiple models on one GPU and fire it up in my unraid machine.
By the time that first bit was done, I realised I could just add one more thing... and another... and well, from beginning of this year when, imho, AI became useful, the whole thing it kind of blew up.
So here's what it does:
- spins a ray cluster
- deploys models using various loaders (vllm, llamacpp, diffusers, sherpa_onnx, etc.)
- serves an openai responses ready api on top of them
Now some the details, already listed in my production readiness (https://docs.model-ship.ai/production-readiness/#production-...) plan:
- scores 17/17 on the Open Responses conformance suite and it's got its own responses redis backed store
- currently supports cpu, metal and cuda
- it has prometheus metrics and grafana dashboard
- can run under docker, k8s or simply native install
- automatically sizes model context
- has MCP support
Feedback would be much appreciated.