A systems walkthrough of vLLM: engine loop, scheduler, paged-attention KV blocks, continuous batching, prefix caching, speculative decoding, disaggregated P/D, multi-GPU executors, a two-node serving stack, and the latency-vs-throughput roofline.
BLOG