A from-scratch C++/CUDA inference engine for five explicitly registered Qwen checkpoints on one NVIDIA GeForce RTX 5090. Startup-frozen residency picks MTP or DFlash speculative decoding, Vision, and one of five KV storage formats; a shared Device KV pool plus pinned Host State/KV checkpoints reuse exact prompt prefixes across 240k-token contexts. Measured aggregate decode reaches 1,146.9 tok/s at concurrency 8 and 15,544.3 tok/s on a 7,680-token prefill.