NVIDIA’s kernel-level serving stack for maximum GPU inference throughput.
Catalog snapshot
NVIDIA-only; the build pipeline is the hard part, not the API.
Fetched 1
TensorRT-LLM is NVIDIA’s serving stack: compiled kernels, in-flight batching, and quantization that push their GPUs to the throughput ceiling. Build-step complexity is real. Apache-based license with NVIDIA terms — read the repo.
TensorRT kernels · in-flight batching · quantization · multi-GPU pipelines