Locality-aware MoE · MiniMax M3

Locality-aware MoE latency on MiniMax M3

FC1 + FC2 latency per rank on Rubin (lower is better). Each tab has its own µs scale.

Mean of 5 repeats, CUDA Graph timing, fixed kernel configurations, HBM at 4752 MHz, balanced routing. avg = geometric mean of the speedups over the 12 token counts shown in the current tab.

Non-localized (baseline) Localized (locality-aware MoE)