NVIDIA's inference optimisation library for LLMs and visual generation models, fully open source since 2025-03-22. Its 28 tech blogs form an auditable engineering ledger: Skip Softmax and sparse attention for long context, the three-part expert parallelism series with one-sided AlltoAll over NVLink, DWDP for NVL72, guided decoding cooperating with speculative decoding, inference-time compute, and evaluating agentic serving with trace replay and job-level metrics. Claims include Llama 4 Maverick above 1,000 TPS per user and 40,000+ tok/s on B200; Bing and NAVER Place run it in production.