High-throughput model server for GPU inference.
Catalog snapshot
Apache 2.0. CUDA is the well-tested path.
Fetched 1
vLLM is how you serve a model to more than one user without writing a batcher. It wants NVIDIA GPUs. Not a chat UI.
PagedAttention · OpenAI API · LoRA · multimodal