Local model runner with a one-line CLI and a REST API.
Catalog snapshot
8 GB RAM is enough for tiny models. 7B class wants more RAM or a GPU.
Fetched 1
Ollama pulls GGUF-class models and serves them on localhost. It is the fastest path from zero to a local Llama. The product is a runner, not a chat UI. Pair it with Open WebUI or a coding agent.
model library · OpenAI-compatible API · GPU auto-detect · Modelfiles