vLLM Reaches 25K Total TPS/GPU on Qwen3.5Aug 6, 2026·9 min readHow vLLM reaches 25K total TPS/GPU on Qwen3.5-397B-A17B-NVFP4 with GB200 NVL72 disaggregated serving, Blackwell GDN kernels, HMA cache transfer, async scheduling fixes, and srt-slurm recipes.