OPEN SOURCE DEEP DIVE
vLLM: the open serving stack that grew out of PagedAttention, with 200+ architectures and five parallelism axes
An inference and serving library out of UC Berkeley's Sky Computing Lab with 2000+ contributors: paged KV memory, continuous batching and chunked prefill, CUDA/HIP graphs with torch.compile, quantisation across FP8/MXFP4/NVFP4/GGUF, pluggable attention kernels (FlashInfer, TRTLLM-GEN, FlashMLA), four speculative decoding routes (n-gram, suffix, EAGLE, DFlash), tensor/pipeline/data/expert/context parallelism with P/D disaggregation, and both OpenAI and Anthropic protocols.
Treating VRAM like virtual memory: the serving stack that grew out of PagedAttention
vLLM started in the Sky Computing Lab at UC Berkeley, and its founding paper is Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023, arXiv:2309.06180). The problem it attacks is concrete: during autoregressive decoding the KV cache size of each request is unknown in advance, so traditional implementations reserve contiguous VRAM for the maximum length. Internal fragmentation, external fragmentation and reservation waste then consume most of the capacity, the number of concurrently resident requests collapses, and throughput collapses with it. PagedAttention borrows paging from the operating system: KV cache is cut into fixed-size blocks, a block table maps logical to physical, and non-contiguous physical blocks can still spell out one request's context. Waste converges to the final partial block.
The point is not "it saved VRAM". The point is that it turned engine scheduling into something that can genuinely be batched: blocks can be shared (prefix caching), swapped out, and migrated between prefill and decode. Nearly everything vLLM does since then grows on that memory model - continuous batching, chunked prefill, prefix caching, P/D disaggregation. Remove the paging layer and each of them has to be redesigned.
Where the throughput comes from: scheduling, graphs and kernels pressed at once
The README lists its performance sources plainly, and they read as three layers:
| Layer | Mechanism | What it removes |
|---|---|---|
| Scheduling | continuous batching, chunked prefill, prefix caching, disaggregated prefill/decode/encode | GPU idle time and redundant compute: a long prefill no longer blocks the whole decode batch, and requests that hit a prefix skip recomputation entirely |
| Execution graph | piecewise and full CUDA/HIP graphs, torch.compile for automatic kernel generation and graph-level transformation | per-step Python and kernel-launch overhead, which dominates at small batch sizes |
| Kernels | attention: FlashAttention, FlashInfer, TRTLLM-GEN, FlashMLA, Triton; GEMM/MoE: CUTLASS, TRTLLM-GEN, CuTeDSL | wasted memory traffic and arithmetic intensity, with grouped GEMM handled separately for MoE |
Pluggable attention kernels deserve their own note. vLLM no longer insists on a single in-house kernel; it wires in FlashInfer, TRTLLM-GEN and FlashMLA and routes by model and hardware. That is an admission that the kernel layer of the inference stack has split into several specialist projects, and that an engine's value now lies in orchestration rather than exclusivity. Speculative decoding likewise offers four routes - n-gram, suffix, EAGLE and DFlash - where the first two cost nothing to train (they guess from context) and the latter two need a draft model or a multi-token prediction head.
Precision and parallelism: making "cheap" a menu
The quantisation list covers essentially every mainstream scheme in play: FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF, compressed-tensors, ModelOpt and TorchAO. GGUF being on that list means llama.cpp's ecosystem of quantised weights can be served directly by a cluster engine - the two previously separate tracks (local inference and datacentre serving) now share a weight format.
Five parallelism axes are exposed: tensor, pipeline, data, expert and context. Expert parallelism is a necessity in the MoE era - models like DeepSeek-V3 carry more experts than any single card can hold, so without EP they cannot be spread across a cluster at all. Context parallelism serves long context by splitting the sequence across cards. Those five knobs plus P/D disaggregation are vLLM's complete answer to rack-scale deployment.
Interfaces and ecosystem: the compatibility layer is wider than the feature layer
The server speaks the OpenAI-compatible API, plus the Anthropic Messages API and gRPC. Supporting two vendor protocols at once is pragmatic: agent frameworks hard-code their clients against one or the other, and each extra protocol an engine speaks removes a layer of migration cost. Structured output runs through xgrammar or guidance, tool calling and reasoning content each get a dedicated parser, and multi-LoRA covers both dense and MoE layers.
More than 200 model architectures are supported, listed by family: decoder-only LLMs (Llama, Qwen, Gemma), MoE LLMs (Mixtral, DeepSeek-V3, Qwen-MoE, GPT-OSS), hybrid attention and state-space models (Mamba, Qwen3.5), multimodal models (LLaVA, Qwen-VL, Pixtral), embedding and retrieval models (E5-Mistral, GTE, ColBERT), and reward and classification models (Qwen-Math). Reward models are not an afterthought - they are the scorer inside the RL post-training loop, which places vLLM as a layer shared by training and inference rather than inference alone.
Hardware reach is equally list-like: NVIDIA, AMD and Intel GPUs plus x86/ARM/PowerPC CPUs are first-class, with plugins for Google TPUs, Intel Gaudi, IBM Spyre, Huawei Ascend, Rebellions NPU, Apple Silicon and MetaX GPUs. Over 2000 contributors from dozens of academic institutions and companies.
Where it sits in the LLM stack
By this site's llm test - "swap it out and either the model's intelligence or its cost per token changes" - vLLM is squarely the second term, and it is the default baseline every other open serving stack gets compared against. Its difference from SGLang is mostly one of emphasis: SGLang leads with RadixAttention prefix reuse and large-scale expert parallelism, while vLLM's strength is architecture coverage, the hardware plugin ecosystem, and its generality as an RL rollout backend. Each keeps re-implementing the other's ideas, and neither has pulled away.
One caveat: "state-of-the-art serving throughput" in the README is the project's own claim with no unified comparison conditions attached. Cross-engine throughput numbers depend heavily on model, parallel configuration, batch size and the input/output length distribution. We label it a project claim; real selection should run benchmarkserving on your own traffic.