Categories
LLM infrastructure.
Serving, gateways, and observability around models. Heavier than a chat UI. Closer to production traffic.
- Rank 1. vLLMHigh-throughput model server for GPU inference.Apache 2.0 · Linux, DockerGitHub stars: 91.6k
- Rank 2. SGLangFast model serving with KV-cache reuse for agent and reasoning workloads.Apache 2.0 · Linux, macOS, +2GitHub stars: 35.9k
- Rank 3. LiteLLMProxy that translates many model APIs into one OpenAI-shaped client.MIT · Linux, macOS, +2GitHub stars: 58.6k
- Rank 4. LangfuseTraces, prompts, and evals for LLM apps.MIT · Linux, Web, +1GitHub stars: 34.5k
- Rank 5. PhoenixLocal LLM tracing and evals workspace from Arize — pip install and open.Other · Linux, macOS, +2GitHub stars: 11.4k
- Rank 6. XinferenceServe LLMs, embeddings, and media models from one Apache-2.0 platform.Apache 2.0 · Linux, macOS, +2GitHub stars: 9.6k
- Rank 7. TensorRT-LLMNVIDIA’s kernel-level serving stack for maximum GPU inference throughput.Other · Linux, Windows, +1GitHub stars: 14.6k
- Rank 8. MLC-LLMCompile-and-serve LLMs on Metal, Vulkan, WebGPU, and phones — no CUDA required.Apache 2.0 · Linux, macOS, +4GitHub stars: 23.1k