OPEN SOURCE DEEP DIVE
llama.cpp: seventeen backends and 1.5-to-8-bit quantisation, taking LLM inference anywhere a compiler exists
An MIT-licensed, dependency-free plain C/C++ inference implementation built on ggml. Seventeen backends span CUDA, Metal, HIP, Vulkan, SYCL, WebGPU, CANN, Hexagon, MUSA, zDNN and ZenDNN; integer quantisation runs 1.5 to 8 bits; CPU+GPU hybrid inference makes models larger than VRAM runnable at all; and GGUF, its output format, is a de facto standard that even vLLM consumes. Two lines - llama cli or llama serve - start an OpenAI-compatible server.
Dependency-free C/C++: moving inference from the datacentre back to the laptop
llama.cpp is an MIT-licensed, plain C/C++ inference implementation for LLMs and VLMs, built on the ggml tensor library. The first line of its own description is "plain C/C++ implementation without any dependencies". That choice costs it the ready-made operator library of the PyTorch ecosystem and buys it the ability to be compiled to almost anywhere a compiler exists - from datacentre GPUs to phones, Raspberry Pi, browsers and IBM mainframes. It is also the de facto source of the GGUF weight format, and vLLM listing GGUF among its supported quantisation formats is an acknowledgement that the format has crossed the project boundary.
Its stated goal is precise: enable LLM (and VLM) inference with minimal setup and state-of-the-art performance across a wide range of hardware, locally and in the cloud. The onboarding path really is four options - guided install via llama.app, Docker, pre-built binaries from the releases page, or clone and build from source. Once installed, two lines are enough:
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF # pull a model straight from Hugging Face and chat
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF # start an OpenAI-compatible API server
Seventeen backends: coverage is the core asset
The backend table is the part of this repository most worth reading line by line, because it doubles as a map of where AI inference can physically happen:
| Backend | Target devices |
|---|---|
| CUDA | NVIDIA GPUs (custom CUDA kernels) |
| Metal | Apple Silicon (first-class citizen, via ARM NEON, Accelerate and Metal) |
| HIP | AMD GPUs |
| Vulkan / OpenCL / WebGPU | general GPUs, Adreno GPUs, browsers and anything with WebGPU |
| SYCL / OpenVINO (in progress) | Intel GPUs, Intel CPUs/GPUs/NPUs |
| CANN | Huawei Ascend NPUs |
| Hexagon | Qualcomm Snapdragon |
| MUSA | Moore Threads GPUs |
| IBM zDNN | IBM Z and LinuxONE mainframes |
| ZenDNN | AMD CPUs |
| BLAS / BLIS / VirtGPU / RPC | generic CPUs, and cross-process / cross-machine work splitting |
CPU instruction-set coverage is just as wide: AVX, AVX2, AVX512 and AMX on x86; RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE on RISC-V. AMX being present means it seriously contemplates matrix multiplication on server CPUs rather than treating the CPU as the GPU's understudy.
Low-bit quantisation and CPU+GPU hybrid inference
Quantisation steps run 1.5, 2, 3, 4, 5, 6 and 8-bit integer. The 1.5-bit tier is one of the most recognisable contributions from the llama.cpp community: mixed precision pushes part of the weights to roughly 1.5 bits per parameter on average, making models that could never reside in consumer memory runnable at all, at the price of visible quality loss. That is the extreme case of trading intelligence for reachability, and it belongs in this site's llm domain precisely because "swap it out and both cost per token and model intelligence change".
Equally important is CPU+GPU hybrid inference: when the model exceeds total VRAM, some layers stay in system memory and are computed on the CPU while the rest go to the GPU. This turns "one consumer card running a model far larger than its VRAM" from impossible into slow-but-usable. It is not a performance feature, it is a reachability feature - it enlarges the set of machines that can run the model rather than making the machines that already could run it faster.
Tooling, and where it sits in the stack
The bundled tools are cli (interactive chat, with VLM sessions), completion, and server (an OpenAI-compatible REST API with a built-in web UI); grammar constraints go through GBNF. Documentation also covers Android builds, multi-GPU usage and a token-generation performance troubleshooting guide.
One positioning point needs to be explicit: llama.cpp is not a cluster serving engine, and it does not compete with vLLM or SGLang for the same job. Those two optimise aggregate throughput and cost per token under concurrent requests; llama.cpp optimises whether a single user on a single machine can get it running with as few dependencies as possible, and how smooth that feels. That is why local developer toolchains (Ollama, LM Studio, Jan and the rest) almost all use it as the substrate, while production serving stacks usually pick vLLM or SGLang. The same model is frequently deployed on both sides: iterate locally on GGUF during development, then switch to a serving engine running FP8 or NVFP4 in production.
Release cadence is continuous nightly - the releases page ships b* tags for nightly builds and v* for formal releases. That is the other difference from academically descended projects: it releases at the pace of a tool, not at the pace of a paper.