Single-file LLM runtime from Mozilla. Download, chmod, run.
Catalog snapshot
Apache 2.0 plus the inner llama.cpp license. CPU works.
Fetched 1
llamafile wraps llama.cpp and a model into one executable. It is the least ops way to try a local model. Not a server farm.
one binary · OpenAI-ish server · GPU optional