Tags
serving.
4 listedUpdated Sep 11, 2026
- Rank 1. vLLMHigh-throughput model server for GPU inference.Apache 2.0 · Linux, DockerGitHub stars: 58k
- Rank 2. SGLangFast model serving with KV-cache reuse for agent and reasoning workloads.Apache 2.0 · Linux, macOS, +2GitHub stars: 36k
- Rank 3. XinferenceServe LLMs, embeddings, and media models from one Apache-2.0 platform.Apache 2.0 · Linux, macOS, +2GitHub stars: 9.6k
- Rank 4. TensorRT-LLMNVIDIA’s kernel-level serving stack for maximum GPU inference throughput.Other · Linux, Windows, +1GitHub stars: 15k