OPEN SOURCE DEEP DIVE
TensorRT-LLM: NVIDIA's specialised-kernel and rack-scale engine, with agentic serving as a first-class topic
NVIDIA's inference optimisation library for LLMs and visual generation models, fully open source since 2025-03-22. Its 28 tech blogs form an auditable engineering ledger: Skip Softmax and sparse attention for long context, the three-part expert parallelism series with one-sided AlltoAll over NVLink, DWDP for NVL72, guided decoding cooperating with speculative decoding, inference-time compute, and evaluating agentic serving with trace replay and job-level metrics. Claims include Llama 4 Maverick above 1,000 TPS per user and 40,000+ tok/s on B200; Bing and NAVER Place run it in production.
The hardware vendor's engine: kernels, a runtime, and a public ledger of tech blogs
TensorRT-LLM is NVIDIA's inference optimisation library for LLMs and visual generation models. It describes itself as three things combined: specialised kernels for common operations, an efficient runtime, and a pythonic framework that lets you customise and extend the system. Since 22 March 2025 it has been fully open source with development moved to GitHub - before that it shipped as a closed component alongside TensorRT, and that history is what separates its old positioning from its current one.
Its structural difference from vLLM and SGLang is that those are hardware-neutral community projects while TensorRT-LLM is the hardware vendor's engine, deeply optimised for its own silicon. Its tech blog therefore reads as a different genre of literature: not "which models do we support" but "how far did we push this operator of this model on this chip".
What it actually optimises, read off the tech blog list
The repository's Tech Blogs list is the most reliable entry point into its engineering priorities, grouped by theme:
| Theme | Post (date) | Problem addressed |
|---|---|---|
| Long-context attention | Accelerating long-context inference with Skip Softmax Attention (02/06); Sparse Attention in TensorRT LLM (03/04) | Attention grows quadratically with sequence length; skipping and sparsification are the two routes back into a servable range |
| MoE and expert parallelism | Scaling Expert Parallelism trilogy (06/05 design and implementation of large-scale EP, 08/01 performance status and optimisation, 10/13 pushing the performance boundary); Optimizing MoE communication with one-sided AlltoAll over NVLink (03/16) | The bottleneck in expert parallelism is all-to-all communication; one-sided communication avoids involving the remote CPU |
| Rack-scale parallelism | DWDP: Distributed Weight Data Parallelism for high-performance LLM inference on NVL72 (04/03) | A rack like NVL72 with 72 cards in one domain needs a new parallelism axis rather than an enlarged old one |
| Disaggregated serving | Disaggregated Serving in TensorRT LLM (06/19) | Split prefill and decode so they scale independently |
| Speculative decoding | N-Gram speculative decoding and auto-enablement (07/26); Combining guided decoding and speculative decoding so CPU and GPU cooperate seamlessly (09/19) | The first costs no draft model and decides automatically whether it is worth enabling; the second fixes structured output and speculative decoding dragging on each other |
| Execution graphs | Tuning CUDA Graph batch sizes for higher output throughput (04/03) | CUDA graphs are bucketed by batch size, and bucket granularity directly sets the replay hit rate |
| Inference-time compute | Inference Time Compute implementation in TensorRT LLM (09/26) | Building "spend more decoding for a better answer" patterns - test-time compute, multi-sample self-consistency - into the engine itself |
| Agentic workloads | Joint optimization of agent applications and TensorRT-LLM (05/15); Evaluating agentic serving with trace replay and job-level metrics (08/20); DeepSeek-V4 on NVIDIA Blackwell: model-specific and agentic-workload optimisations (07/17) | Agent traffic looks nothing like chat: multi-turn, heavy prefix reuse, tool calls, and measured per job rather than per request |
| Visual generation | Scaling video generation across the NVL72 rack (07/01); Accelerating video generation with GEMM quantisation, attention quantisation and Skip Softmax (09/02) | Diffusion models enter the same engine; visual generation diffusion is officially supported from 2026-04-03 |
Two threads deserve separate mention. First, agentic serving is now a first-class topic: not only joint optimisation but a proposed evaluation method - trace replay plus job-level metrics - which is an admission that measuring agent load by per-request QPS and latency is simply wrong. Second, diffusion models have entered the same engine, the same direction as SGLang Diffusion, meaning "LLM inference engine" is expanding into "generative model inference engine".
The published performance claims
Quantified results available from the news history include: Blackwell breaking the 1,000 TPS/user barrier with Meta's Llama 4 Maverick; Llama 4 running at over 40,000 tokens per second on B200 GPUs; a 3x inference throughput gain on Llama 3.3 70B from speculative decoding (with a separate 3.6x overall throughput claim); and an NVIDIA Blackwell world record for DeepSeek-R1 inference performance, with nvidia/DeepSeek-R1-FP4 weights published on Hugging Face. Day-0 support covers OpenAI's GPT-OSS-120B and 20B, and LG AI Research's EXAONE 4.0. Named production adopters include Bing (its transition to LLM/SLM models for search) and NAVER Place (SLM-based vertical services).
Every one of those numbers comes from NVIDIA's own blog or release material, with no independent third-party reproduction on the same hardware, so we record them as vendor claims. FP4 is the item worth attention: it is Blackwell's native low-precision format, compressing weights and activations to the 4-bit range, and cost per token changes step-wise at that tier - while it is currently available mainly on this vendor's hardware.
KV cache reuse, and where it sits in the stack
In January 2025 it published KV cache reuse optimisations on its own. That is the same problem solved at three different layers by three projects: block-level reuse inside the engine here, prefix matching via a radix tree in SGLang, and cache-overlap-aware routing at cluster level in Dynamo. The three stack; they are not mutually exclusive options.
The practical selection conclusion: if the deployment target is NVIDIA hardware and you want the FP4 and NVL72 rack-scale parallelism dividends, TensorRT-LLM's kernel depth is not something community engines close quickly. If you need cross-hardware reach (AMD, Intel, TPU, Ascend) or the widest model-architecture compatibility, vLLM and SGLang cost less effort. Dynamo accepting all three as backends is the vendor's own admission that this choice is workload-dependent rather than globally unique.