Serve LLMs, embeddings, and media models from one Apache-2.0 platform.
Catalog snapshot
One-box alternative to juggling Ollama plus separate embedding servers.
Fetched 1
Xinference launches, schedules, and serves many models from one platform: LLMs, embeddings, rerankers, image and audio models behind an OpenAI-shaped API with a built-in chat UI. Apache-2.0 and very active.
multi-model serving · built-in UI · GPU scheduling · OpenAI-compatible API