Kimi K3 Performance Optimizations in vLLM: The Road to 2.8× Throughput
·7 min read
Kimi K3 serving optimizations across scheduling, KDA prefix caching, ReplaySSM state recovery, PD disaggregation and state offload, parallelism, MoE, and GPU kernels.
3 posts
Kimi K3 serving optimizations across scheduling, KDA prefix caching, ReplaySSM state recovery, PD disaggregation and state offload, parallelism, MoE, and GPU kernels.

vLLM delivers day-0 Kimi K3 serving with hybrid KDA prefix caching, DSpark speculative decoding, production-scale disaggregation, and optimized kernels across NVIDIA and AMD GPUs.

A preview of production-scale Kimi K3 support in vLLM, including KDA-aware prefix caching, fused kernels, optimized MXFP4 MoE, multimodal integration, and initial NVIDIA and AMD paths.