vime × RL-Kernel × AMD: Bitwise Train–Rollout Consistency on ROCm

10 min read
RL-Kernel Team, vime Team, and AMD Team
  • vime connects Megatron training, vLLM rollout, and a Data Buffer into a complete RL post-training workflow. RL-Kernel makes Megatron and vLLM follow the same numerical execution contract when computing logprobs.
  • In an end-to-end Qwen3-8B GRPO experiment on AMD Instinct MI300X, the vime + RL-Kernel strict path ran for 200 consecutive steps with mismatch_count = 0 and max_abs_diff = 0 throughout.
  • This post explains why train–rollout mismatches occur, how vime and RL-Kernel divide the work, and how we implemented and validated bitwise consistency while preserving native ROCm execution paths.

Why RL Training Needs Train–Rollout Numerical Consistency

Training and rollout usually use different execution engines and operators. Even with the same model, weights, and inputs, differences in kernels, parallelism, and reduction order can produce different logprobs.

This is a long-standing issue in RL systems because it affects the importance ratio, rollout–training KL diagnostics, and clipping. Existing work has focused mainly on NVIDIA platforms, while systematic exploration on ROCm has been more limited.

Building on vime × RL-Kernel, this work aligns the Attention, FFN, logprob, and communication paths on AMD Instinct MI300X to achieve bitwise train–rollout consistency.

Concretely, the rollout engine generates token from prefix and records the rollout-side logprob . The training engine then recomputes the same token under the same weight version to obtain . We track the difference and the corresponding importance ratio .

Before the policy update, both sides should be scoring the same token with the same weights. The ideal target is therefore and . Any non-zero difference introduces an additional numerical policy shift, which can propagate into KL and clipping. The strict target of this integration is to make the logprob difference bitwise zero.

Why the Same Model Can Produce Different Results

GPUs use finite-precision floating-point arithmetic, so rounding can occur at every step. Consider a simple example with , , and . Computing gives 1, while may give 0 because can round back to at finite precision.

Training and inference encounter the same issue. Training is optimized for packed sequences, backpropagation, and multi-GPU parallelism; inference is optimized for prefill, decode, dynamic batching, and paged KV cache. Even when the model, weights, and inputs are identical, the two sides may use different block sizes, Split-K or Split-KV strategies, reduction orders, fusion patterns, and intermediate precision.

Therefore, the same model does not necessarily execute the same floating-point computation in training and rollout, nor does it guarantee identical logprobs.

We describe the observable numerical behavior of a compute node with a numerical execution contract:

  • : the elements on which the output actually depends, together with the reduction domain.
  • : how the reduction domain is partitioned into partial results.
  • : the order in which partial results are merged.
  • : the precision of inputs, accumulators, intermediate states, and outputs.
  • : where rounding or downcasting occurs.
  • : the numerical primitives actually used, such as exp, log, rsqrt, SiLU, and FMA.

Bitwise consistency requires training and rollout to follow the same observable numerical contract along the forward path being compared.

vime Aligns the Timeline; RL-Kernel Aligns Numerical Execution

Their respective roles are:

  • vime aligns the training timeline: which token batch belongs to which step, which weight version generated it, which rollout record enters an update, and when new weights are synchronized to vLLM.
  • RL-Kernel aligns numerical execution: which values participate in a computation, how they are partitioned and merged, which intermediate precision is used, and where rounding occurs.

This division of labor is the same on CUDA and ROCm, but the numerical rules must ultimately be implemented in each platform's operators, compiler, and communication stack. Paths already validated on CUDA therefore had to be adapted and revalidated on ROCm. We added deterministic MFMA-based GEMM kernels, vocabulary reduction, and HIP IPC communication; fixed the execution schedule of AITER/CK Attention; addressed last-bit differences caused by math functions and compiler fusion; and fixed state issues in paged KV layout and HIP Graph replay.

After these adaptations, the weights and tokens aligned by vime could pass through training and inference along the same numerical path. On an 8× MI300X Qwen3-8B configuration, the resulting logprobs remained bitwise identical for 200 consecutive training and rollout steps.

Confirming That Both Sides Compute the Same Object

Before comparing floating-point results, we verify that the following match:

  • checkpoint and weight version;
  • prefix, token, and active mask;
  • position, RoPE, causal mask, and padding mask;
  • logical K/V after paged KV cache mapping;
  • global sequence/head/vocabulary indexing and shard-to-global mapping;
  • the true vocabulary range and any random state relevant to the comparison.

If any condition differs, the sample is marked , and the final difference is not attributed to kernels.

Why Transformer Numerical Divergences Must Be Handled Together

RMSNorm, GEMM, Attention, linear logprob, and distributed collectives look like separate modules, but all of them merge partial results. Block size, Split-K or Split-KV, reduction trees, collective trees, exp and log implementations, intermediate precision, and fusion can all change the merge order and rounding boundaries.

Attention's Split-KV merge, linear logprob's cross-TP vocabulary merge, GEMM's K-dimension reduction, and ROCm collectives are therefore parts of one numerical chain. The RL-Kernel strict path fixes these boundaries together; pinning down a single kernel or collective is not enough to guarantee bitwise-identical logprobs.

Which Boundaries RL-Kernel Fixes in vime

The RL-Kernel strict path covers the main forward boundaries that determine logprobs:

Compute boundaryWhat the strict path fixes
RMSNormReduction domain, epsilon, residual addition, and output boundary
AttentionPosition, mask, logical paged KV, split policy, LSE precision, and final cast
GEMM and SwiGLUK-dimension reduction, accumulation precision, epilogue, activation, and materialization boundary
Linear logprobTrue vocabulary range, target ownership, local reduction, and cross-rank LSE merge
Distributed collectivesPayload ownership, dtype, and a fixed rank reduction order

This contract does not require training and rollout to share every memory layout or scheduling policy. The two engines can still optimize independently as long as those optimizations do not change the numerical semantics of the compared results.

200-Step Alignment Experiment with vime on ROCm

We completed a strict 200-step validation on ROCm within the full vime workflow of Megatron training and vLLM rollout, with zero mismatches throughout. vime handled rollout, training, weight synchronization, and the sample lifecycle, while RL-Kernel aligned the numerical execution path used to compute logprobs in the same pipeline.

Experimental Setup

ItemConfiguration
Model / dtypeQwen3-8B / BF16
Hardware1 node, 8× AMD Instinct MI300X 192GB
MegatronTP4 / CP2 / PP1, using 8 GPUs
Rollout2 vLLM engines, TP4 each
PlacementActor and rollout colocated
Horizon200 rollout/training steps
Datasetdapo-math-17k
SeedsTraining 1234, rollout 1234
Sampling1 prompt × 8 samples per step, global batch 8
Response limit6,912 tokens
Dynamic batchingMaximum 4,096 tokens/GPU
vLLM memory utilization0.38
HIP GraphFULL_AND_PIECEWISE, preserving the production graph execution path
KL lossEnabled, coefficient 0.001
ValidationInputs and source values were frozen before and after the run; every step had to pass provenance and mismatch checks

Across all 200 steps of the strict path, both mismatch_count and max_abs_diff remained zero.

200-Step Training Trajectory

Figure 1 plots the train–rollout mismatch count and maximum absolute on the same 200-step timeline. The RL-Kernel strict path remains at zero throughout, while native vime shows a mismatch at every step.

Consistency comparison between native vime and vime plus RL-Kernel
Consistency comparison between native vime and vime plus RL-Kernel

Figure 1: Consistency comparison between native vime and vime + RL-Kernel.

These signals appear over the same interval and are consistent with persistent train–rollout mismatch. Together, they provide end-to-end evidence for strict alignment. The experiment shows that vime + RL-Kernel can maintain verifiable bitwise consistency and a more stable training trajectory throughout the 200-step run.

Figure 2 shows the mean absolute train–rollout logprob difference over 200 steps. vime + RL-Kernel remains at zero throughout.

Mean absolute train-rollout logprob difference across 200 ROCm steps
Mean absolute train-rollout logprob difference across 200 ROCm steps

Figure 2: Mean absolute train–rollout logprob difference across 200 steps.

What vime × RL-Kernel Achieves on ROCm

  • Bitwise consistency: Training and rollout logprobs match exactly on ROCm. Across all 200 steps, mismatch_count remains zero and the maximum logprob difference is also zero.
  • Stable consistency guarantees: Zero mismatch is maintained throughout the 200-step end-to-end training run, making results easier to verify and reproduce.
  • Complete ROCm execution evidence: The validation records the kernels, HIP Graph execution, paged KV, collectives, and fallback paths actually used at runtime.
  • Fast failure localization: Operator ablations identify the specific operator or system boundary where train–rollout divergence begins.

On 8× AMD Instinct MI300X, RL-Kernel + vime maintained zero mismatch across all 200 steps.

Current Scope and Next Steps

The current end-to-end validation covers Qwen3-8B Dense, vime, vLLM, Megatron-LM, and AMD Instinct MI300X. Next, we plan to extend the work to more MoE and multimodal models and additional AMD GPU architectures, while continuing to optimize the ROCm strict path.

Acknowledgments

This integration of vime, RL-Kernel, and AMD was made possible by the support of our partner organizations and the open-source community.

  • AMD: Liz Li and Yuhan Yang, for providing AMD Instinct GPU compute resources, in-depth technical collaboration, and long-term support that enabled the end-to-end validation of vime + RL-Kernel on ROCm.

  • Inferact: Ao Shen, vime maintainer, for his trust and support in the vime integration, community coordination, and ongoing maintenance.

  • Moore Threads: Lei Ding, for advancing RL-Kernel's MUSA support.

  • Huawei: Yang Chen, for advancing RL-Kernel's Ascend support.

  • Embedded LLM: For supporting the project's development and community collaboration.

  • vLLM community: For its close collaboration with RL-Kernel.

  • RL-Kernel v0.1.0 core contributors: Chutian Wang, Jiajie Li, Siru He, Xiaosong Ma, Kaijie Lin, Jian Zhang, Huihong Lu, Yunxiang Cai, Bosong Yang, Zhewei Liu, Houhong Liang, Ryan Huang, and Vensen Mu.

  • Community contributors whose PRs were merged into v0.1.0: Xiaopeng Du, Yuepeng Pan, Yiyang Fei, Ziying Tao, Zhifu Liu, Zhengtao Chen, Mengjie Li, Zien Liu, and GitHub users haoruilee, luoyueyuguang, hongleng, and smarslou.

The implementation is open source in the RL-Kernel repository. We welcome feedback and discussion through GitHub Issues and Pull Requests, as well as contributions that extend train–rollout consistency support to more models, hardware platforms, and RL post-training workloads.