Skip to content
Alternatives

OpenAI API.

Local OpenAI-compatible servers, gateways, and fine-tuning stacks for your own models.

26 listedCatalog Sep 1, 2026

  1. Rank 1. llamafileSingle-file LLM runtime from Mozilla. Download, chmod, run.GitHub stars: 23k
  2. Rank 2. LiteLLMProxy that translates many model APIs into one OpenAI-shaped client.GitHub stars: 30k
  3. Rank 3. vLLMHigh-throughput model server for GPU inference.GitHub stars: 58k
  4. Rank 4. llama.cppC/C++ runtime that made local GGUF models practical.GitHub stars: 87k
  5. Rank 5. exoLink everyday devices into one cluster and run models too big for any single machine.GitHub stars: 47k
  6. Rank 6. shell_gptCLI that runs an LLM as a shell command for scripts, refactors, and answers.GitHub stars: 12k
  7. Rank 7. DifyVisual LLM app builder with RAG, agents, and a backend.GitHub stars: 118k
  8. Rank 8. MLXApple’s framework for running and training models on Mac unified memory.GitHub stars: 28k
  9. Rank 9. LangfuseTraces, prompts, and evals for LLM apps.GitHub stars: 16k
  10. Rank 10. OllamaLocal model runner with a one-line CLI and a REST API.GitHub stars: 152k
  11. Rank 11. LLaMA-FactoryWeb UI and YAML recipes for fine-tuning many open LLMs.GitHub stars: 58k
  12. Rank 12. torchtunePyTorch-native library for fine-tuning LLMs with recipes.GitHub stars: 7k
  13. Rank 13. DeepSpeedMicrosoft’s Apache-2.0 ZeRO optimizer for multi-GPU, multi-node LLM training.GitHub stars: 43k
  14. Rank 14. PhoenixLocal LLM tracing and evals workspace from Arize — pip install and open.GitHub stars: 11k
  15. Rank 15. LocalAIOpenAI-compatible API that can sit in front of several local backends.GitHub stars: 35k
  16. Rank 16. LlamaIndexThe MIT Python framework for building document ingestion and RAG pipelines.GitHub stars: 52k
  17. Rank 17. UnslothFaster LoRA fine-tuning with lower VRAM on consumer GPUs.GitHub stars: 45k
  18. Rank 18. TRLHugging Face’s Apache-2.0 trainers for SFT, DPO, PPO, and reward modeling.GitHub stars: 19k
  19. Rank 19. nanoGPTKarpathy’s ~300-line GPT trainer — the cleanest way to learn LLM training.GitHub stars: 63k
  20. Rank 20. SGLangFast model serving with KV-cache reuse for agent and reasoning workloads.GitHub stars: 36k
  21. Rank 21. XinferenceServe LLMs, embeddings, and media models from one Apache-2.0 platform.GitHub stars: 9.6k
  22. Rank 22. MLC-LLMCompile-and-serve LLMs on Metal, Vulkan, WebGPU, and phones — no CUDA required.GitHub stars: 23k
  23. Rank 23. GraphRAGMicrosoft’s MIT library for graph-based RAG over whole document corpora.GitHub stars: 36k
  24. Rank 24. AxolotlYAML-driven fine-tuning toolkit for serious training runs.GitHub stars: 11k
  25. Rank 25. PEFTHugging Face’s LoRA/QLoRA library that makes big-model tuning fit one GPU.GitHub stars: 22k
  26. Rank 26. TensorRT-LLMNVIDIA’s kernel-level serving stack for maximum GPU inference throughput.GitHub stars: 15k