DSv4-Pro hybrid KV cache
MXFP4 C4I
FP8 C4I
KVCacheManager
→
G1
C4A + C4I + C128A
G2
Sliding window cache · 31
G3
Sliding window cache · 30
G4
C128 compressor state
G5
C4A/C4I compressor state
token-proportional
per-request
→
Shared BlockPool
Before · size-bucketed tensors
31 large + 30 index + 31 small = 92 tensors
layer slot 1
layer slot 2
…
G1
G2
G3
G4
G5
After · packed by cache group
one physical backing tensor
layer pages concatenated left-to-right →
block stride
G1
G2
G3
G4
G5
C4A cache
C4I cache
C128A cache
Sliding window cache
C128 compressor state
C4I compressor state
C4A compressor state
padding waste
idle / wasted