OPEN SOURCE DEEP DIVE
Dynamo: the datacenter-scale orchestration layer above inference engines, with KV-aware routing, tiered KVBM and an SLA-driven planner
NVIDIA's open-source (Apache-2.0) orchestration layer, Rust for the performance path and Python for extensibility. It explicitly does not replace SGLang, TensorRT-LLM or vLLM - it wires them into a coordinated multi-node system: disaggregated prefill/decode, KV-aware routing (2x faster TTFT on Qwen3-Coder 480B), KVBM offloading KV cache to CPU/SSD/remote, ModelExpress streaming weights over NVLink for 7x faster cold starts, an SLA-driven Planner (80% fewer breaches at 5% lower TCO in Alibaba production), and Grove for topology-aware NVL72 scheduling.
The layer above the engines: turning a pile of GPUs into one inference system
Dynamo is NVIDIA's open-source (Apache-2.0), datacenter-scale inference stack, and the first line of its README states its job precisely: it is the orchestration layer above inference engines - it does not replace SGLang, TensorRT-LLM or vLLM, it turns them into a coordinated multi-node inference system. Performance-critical paths are written in Rust; the extension surface is Python.
That positioning matters because it makes Dynamo a layer above those projects rather than a competitor to them. A single engine optimises "how fast can one card or one node emit tokens"; Dynamo optimises "given a cluster of nodes that are each already fast, how do they stop wasting each other's work". It also states the boundary plainly: if you are running a single model on a single GPU, your inference engine alone is probably sufficient.
When you need it: five tests
The README's list of intended scenarios reads directly as a selection checklist: serving LLMs across multiple GPUs or nodes and needing to coordinate them; wanting KV-aware routing to avoid redundant prefill computation; needing to scale prefill and decode independently (disaggregated serving); wanting automatic scaling that meets latency SLAs at minimum total cost of ownership; and needing fast cold starts when spinning up new replicas. If one of the five does not hold, start with the engine alone.
Six core components
| Component | What it does | Why it matters |
|---|---|---|
| Disaggregated prefill/decode | Separates prefill and decode into independently scalable GPU pools | The two phases have different compute profiles (prefill is arithmetic-bound, decode is memory-bound); only split apart can each run on hardware tuned for it, and GPU utilisation rises |
| KV-aware routing | Routes requests by worker load and KV cache overlap | A prefix that already exists does not need prefill again - the published result is 2x faster TTFT |
| KVBM (KV Block Manager) | Offloads KV cache across GPU → CPU → SSD → remote storage | Effective context length exceeds VRAM; for long context and long agentic sessions the capacity problem becomes tiered storage instead of buying more cards |
| ModelExpress | Streams model weights GPU-to-GPU over NIXL/NVLink | 7x faster cold start for new replicas - and scale-up speed is what determines whether autoscaling can actually keep up with traffic |
| Planner | SLA-driven autoscaler that profiles the workload and right-sizes each pool | Meets latency targets at minimum TCO, turning "how much headroom do we keep" from operator folklore into something computable |
| Grove | Kubernetes operator for topology-aware gang scheduling (NVL72) | At rack scale, which rack, host and NUMA node a replica lands on directly changes its communication cost |
All six answer the same question: once inference stops being "one process on one machine" and becomes "a set of services across a datacentre", who absorbs the extra cost - redundant prefill, cold starts, autoscaling headroom, topology mismatch. Dynamo's answer is the orchestration layer.
Published results, and how to read them
| Result | Context |
|---|---|
| 7x higher throughput per GPU | DeepSeek R1 on GB200 NVL72 with Dynamo versus B200 without (SemiAnalysis InferenceX) |
| 750x higher throughput | DeepSeek-R1 on GB300 NVL72 (InferenceX v2) |
| 7x faster model startup | ModelExpress weight streaming, DeepSeek-V3 on H200 |
| 2x faster time to first token | KV-aware routing, Qwen3-Coder 480B (Baseten benchmark) |
| 80% fewer SLA breaches | Planner autoscaling at 5% lower TCO (Alibaba APSARA 2025) |
Two cautions apply to this table. First, the "7x" row changes hardware (GB200 NVL72 versus B200) and software (with Dynamo versus without) at the same time, so it compares a whole rack solution against a single card rather than isolating software gain; the "750x" row spans two rack generations for the same reason. Second, the 2x TTFT and 80% SLA rows come from third-party settings - Baseten and Alibaba Cloud - and are more trustworthy than NVIDIA's own comparisons because they name a concrete workload (Qwen3-Coder 480B) and a concrete metric. We record the first two as project claims and attribute the third-party two to their sources.
Backend support and governance
The feature matrix is broken out per backend: disaggregated serving, KV-aware routing, the SLA planner, multimodal and tool calling are all green across SGLang, TensorRT-LLM and vLLM; KVBM is marked in-progress on the SGLang side and supported on the other two. The full matrix also covers LoRA, request migration, speculative decoding, and feature interactions. Accepting all three engines as backends is itself the evidence that the orchestration layer does not pick sides.
Two governance details are worth recording. The repository manages design proposals through a full label lifecycle - dep:draft / proposed / approved / implementing / completed / deferred / superseeded - so design debate is public and traceable. And it maintains a public community calendar with regular meetings (for example the Baseten × Dynamo × SGLang RL post-training meetup on 2026-09-10), which says that serving RL post-training is already a formal agenda item on this stack rather than an adjunct to inference serving.
Reading it against Neroued/ninfer, already in this index, is instructive: ninfer narrows scope to the extreme (one card, five registered weight sets, rejecting any architecture other than sm_120a) to buy single-card peak throughput, while Dynamo widens scope to the extreme (across racks, engines and storage tiers) to buy cluster efficiency. Both ends of the inference stack have serious work happening; the middle - general serving on a single multi-GPU node - is where vLLM and SGLang actually fight.