C/C++ runtime that made local GGUF models practical.
Catalog snapshot
CPU inference works. It is slow on 7B+ without a GPU.
Fetched 1
llama.cpp is the engine under a lot of local AI. It is a CLI and a library, not a product UI. If you need to squeeze a 7B model onto 8 GB, this is the stack to learn.
GGUF · CPU and GPU backends · server example · quantization