Locality-aware MoE · MiniMax M3
FC1 + FC2 latency per rank on Rubin (lower is better). Each tab has its own µs scale.
Mean of 5 repeats, CUDA Graph timing, fixed kernel configurations, HBM at 4752 MHz, balanced routing. avg = geometric mean of the speedups over the 12 token counts shown in the current tab.