An inference and serving library out of UC Berkeley's Sky Computing Lab with 2000+ contributors: paged KV memory, continuous batching and chunked prefill, CUDA/HIP graphs with torch.compile, quantisation across FP8/MXFP4/NVFP4/GGUF, pluggable attention kernels (FlashInfer, TRTLLM-GEN, FlashMLA), four speculative decoding routes (n-gram, suffix, EAGLE, DFlash), tensor/pipeline/data/expert/context parallelism with P/D disaggregation, and both OpenAI and Anthropic protocols.
A high-performance serving framework hosted by LMSYS: RadixAttention prefix caching, a zero-overhead CPU scheduler, prefill-decode disaggregation, DFlash and Spec V2 speculative decoding, compressed finite state machines for structured output, and large-scale expert parallelism (96 H100 GPUs; 3.8x prefill and 4.8x decode on GB200 NVL72 part II). Its day-0 ledger covers Kimi K3, DeepSeek-V4, GLM5.2 NVFP4, Nemotron 3 and MiniMax M2, and AReaL, Miles, slime, Tunix and verl all use it as an RL rollout backend.
NVIDIA's inference optimisation library for LLMs and visual generation models, fully open source since 2025-03-22. Its 28 tech blogs form an auditable engineering ledger: Skip Softmax and sparse attention for long context, the three-part expert parallelism series with one-sided AlltoAll over NVLink, DWDP for NVL72, guided decoding cooperating with speculative decoding, inference-time compute, and evaluating agentic serving with trace replay and job-level metrics. Claims include Llama 4 Maverick above 1,000 TPS per user and 40,000+ tok/s on B200; Bing and NAVER Place run it in production.