Prefill of one n-token prompt; schematic, not to scale.
Layer ids from the DeepSeek-V4.1-Flash config: the last KV-source layer is 20. Decoder-side replay: #58132; CUDA graphs for the trimmed layers: #59532.
computed on all tokensglobal KV (layer 20)computed on the last 128 tokensskipped