
DeepSeek-V4.1-Flash on vLLM: 5x Agentic Throughput Since Day 0
·12 min read
Within three weeks of release, vLLM made DeepSeek-V4.1-Flash 1.9x faster at low concurrency and lifted its throughput 5x on SemiAnalysis AgentX, with SWA bounded replay, CUDA graphs, DeepSeek's new kernels, and vLLM kernel fusions.