Fast model serving with KV-cache reuse for agent and reasoning workloads.
Catalog snapshot
Shines where prompt prefixes repeat; pairs well with agent traffic.
Fetched 1
SGLang serves models with RadixAttention — KV-cache reuse across requests — which makes it unusually fast for agents and reasoning loops that replay the same prefixes. Apache-2.0, OpenAI-compatible endpoints, and a very active repo.
RadixAttention KV cache · structured generation · OpenAI API · multi-LoRA serving