GPUStack
FreeAn open-source GPU cluster manager for running LLMs
About GPUStack
GPUStack is an open-source GPU cluster manager designed for high-performance AI model serving and on-demand GPU instance provisioning. It orchestrates inference engines like vLLM, SGLang, and TensorRT-LLM, supports custom engines, and provides SSH-accessible GPU instances for development, fine-tuning, and interactive workloads. Key capabilities include multi-cluster management across on-premises, Kubernetes, and cloud environments, day-0 model support with pluggable engine architecture, performance-optimized configurations (low latency, high throughput, LMCache, HiCache, speculative decoding), and enterprise-grade operations with automated failure recovery, load balancing, monitoring, authentication, and access control. GPUStack supports a wide range of accelerators including NVIDIA, AMD, Ascend NPU, Hygon DCU, and MThreads GPUs, and offers industry-standard APIs for LLM, voice, image, and video models.
Key Features
Pros & Cons
- Open-source with no licensing fees
- Supports a broad range of accelerators beyond NVIDIA GPUs
- Performance-optimized configurations deliver strong inference performance out of the box
- Enterprise-grade features like monitoring, authentication, and failure recovery
- Multi-cluster management enables unified GPU orchestration across different environments
- Relatively new project with a smaller community compared to established cluster managers
- May require significant infrastructure and expertise to set up and maintain
- Documentation and tutorials are still evolving