Alternatives
OpenAI API.
Local OpenAI-compatible servers, gateways, and fine-tuning stacks for your own models.
26 listedCatalog Sep 1, 2026
- Rank 1. llamafileSingle-file LLM runtime from Mozilla. Download, chmod, run.Apache 2.0 · Linux, macOS, +1GitHub stars: 23k
- Rank 2. LiteLLMProxy that translates many model APIs into one OpenAI-shaped client.MIT · Linux, macOS, +2GitHub stars: 30k
- Rank 3. vLLMHigh-throughput model server for GPU inference.Apache 2.0 · Linux, DockerGitHub stars: 58k
- Rank 4. llama.cppC/C++ runtime that made local GGUF models practical.MIT · Linux, macOS, +1GitHub stars: 87k
- Rank 5. exoLink everyday devices into one cluster and run models too big for any single machine.Apache 2.0 · Linux, macOSGitHub stars: 47k
- Rank 6. shell_gptCLI that runs an LLM as a shell command for scripts, refactors, and answers.MIT · Linux, macOS, +1GitHub stars: 12k
- Rank 7. DifyVisual LLM app builder with RAG, agents, and a backend.Other · Linux, Web, +1GitHub stars: 118k
- Rank 8. MLXApple’s framework for running and training models on Mac unified memory.MIT · macOSGitHub stars: 28k
- Rank 9. LangfuseTraces, prompts, and evals for LLM apps.MIT · Linux, Web, +1GitHub stars: 16k
- Rank 10. OllamaLocal model runner with a one-line CLI and a REST API.MIT · Linux, macOS, +2GitHub stars: 152k
- Rank 11. LLaMA-FactoryWeb UI and YAML recipes for fine-tuning many open LLMs.Apache 2.0 · Linux, Windows, +1GitHub stars: 58k
- Rank 12. torchtunePyTorch-native library for fine-tuning LLMs with recipes.BSD 3-Clause · LinuxGitHub stars: 7k
- Rank 13. DeepSpeedMicrosoft’s Apache-2.0 ZeRO optimizer for multi-GPU, multi-node LLM training.Apache 2.0 · Linux, macOS, +2GitHub stars: 43k
- Rank 14. PhoenixLocal LLM tracing and evals workspace from Arize — pip install and open.Other · Linux, macOS, +2GitHub stars: 11k
- Rank 15. LocalAIOpenAI-compatible API that can sit in front of several local backends.MIT · Linux, macOS, +2GitHub stars: 35k
- Rank 16. LlamaIndexThe MIT Python framework for building document ingestion and RAG pipelines.MIT · Linux, macOS, +1GitHub stars: 52k
- Rank 17. UnslothFaster LoRA fine-tuning with lower VRAM on consumer GPUs.Apache 2.0 · Linux, Windows, +1GitHub stars: 45k
- Rank 18. TRLHugging Face’s Apache-2.0 trainers for SFT, DPO, PPO, and reward modeling.Apache 2.0 · Linux, macOS, +1GitHub stars: 19k
- Rank 19. nanoGPTKarpathy’s ~300-line GPT trainer — the cleanest way to learn LLM training.MIT · Linux, macOS, +1GitHub stars: 63k
- Rank 20. SGLangFast model serving with KV-cache reuse for agent and reasoning workloads.Apache 2.0 · Linux, macOS, +2GitHub stars: 36k
- Rank 21. XinferenceServe LLMs, embeddings, and media models from one Apache-2.0 platform.Apache 2.0 · Linux, macOS, +2GitHub stars: 9.6k
- Rank 22. MLC-LLMCompile-and-serve LLMs on Metal, Vulkan, WebGPU, and phones — no CUDA required.Apache 2.0 · Linux, macOS, +4GitHub stars: 23k
- Rank 23. GraphRAGMicrosoft’s MIT library for graph-based RAG over whole document corpora.MIT · Linux, macOS, +1GitHub stars: 36k
- Rank 24. AxolotlYAML-driven fine-tuning toolkit for serious training runs.Apache 2.0 · Linux, DockerGitHub stars: 11k
- Rank 25. PEFTHugging Face’s LoRA/QLoRA library that makes big-model tuning fit one GPU.Apache 2.0 · Linux, macOS, +1GitHub stars: 22k
- Rank 26. TensorRT-LLMNVIDIA’s kernel-level serving stack for maximum GPU inference throughput.Other · Linux, Windows, +1GitHub stars: 15k