Compile-and-serve LLMs on Metal, Vulkan, WebGPU, and phones — no CUDA required.
Catalog snapshot
Runs chat models on phones and in browsers — the deployment surface is the point.
Fetched 1
MLC-LLM compiles models to run anywhere ML compilation reaches: Apple Silicon, AMD, Android, iOS, and even WebGPU in the browser, with an OpenAI-compatible REST server. Apache-2.0. The escape hatch from CUDA-only serving.
ML compilation · Metal, Vulkan, and WebGPU backends · REST server · iOS and Android targets