AI Pulse

LLM后训练选RL算法,用旧数据提吞吐,可能让更新噪声变大

LLM后训练选RL算法,用旧数据提吞吐,可能让更新噪声变大

Choosing an RL algorithm for LLM post-training means sorting through a growing list of options: PPO, GRPO, Dr.GRPO, GSPO, DAPO, CISPO. The choice also depends on how you generate rollouts and when you train on them. An update that works well with fresh samples may behave differently when generation and training run asynchronously.

These decisions become more important as RL takes on longer tasks. An agent might spend minutes or hours writing code, calling tools, and checking its work before receiving a reward. Completion times can vary widely, so waiting for the slowest trajectories leaves hardware idle, but continuing to train means the policy changes while those trajectories are still running.

This creates a practical tradeoff. Training on older rollouts can keep hardware busy, but correcting for the mismatch can make updates noisy, especially over long sequences. Limiting those corrections can stabilize training while introducing bias. The question is whether the extra throughput outweighs the loss in update quality.

The goal is to improve the model as efficiently and stably as possible within a particular compute budget. The value of a faster pipeline depends on the bias and variance of the updates it produces. We can use bias, variance, and throughput to compare these choices and understand their tradeoffs.

更新中的偏差与方差

First, what are we actually estimating and optimizing for? The model’s policy πθ assigns probabilities to possible responses. We sample a prompt x from a distribution D, generate a response y, and score it with a reward R(x, y). Our starting objective is expected reward:
J(\theta)=\mathbb E_{x\sim\mathcal D,\;y\sim\pi_\theta(\cdot\mid x)}[R(x,y)].

The policy gradient points in the direction of increasing expected reward. We estimate it from sampled responses. When those responses come from the current policy (on-policy sampling), the reward does not depend directly on the model’s parameters, and there is no KL penalty, a basic estimator is:
\hat g=(R(x,y)-b(x))\sum_{t=1}^{T}\nabla_\theta\log\pi_\theta(y_t\mid x,y_{<t}).
Here T includes the stopping decision, and b(x) is a baseline, a reference reward for the prompt. The update reinforces responses that outperform that reference and discourages those that underperform it.

For a given prompt, subtracting a baseline that is independent of the sampled response leaves the expected gradient unchanged. A well-chosen baseline makes the sampled updates less variable around that expectation. It can therefore reduce variance without introducing bias. [2]

For this gradient estimate, bias is a systematic difference from the target gradient, while variance measures how much the estimate fluctuates across samples. Several sources of bias matter in practice:

Estimator bias: The expected policy-gradient estimate differs from the gradient of the chosen objective, as when importance weights are capped.
Objective bias: We change the objective or relative sample weights, through choices such as length normalization or KL regularization.
Selection bias: Admission rules change which data reaches training, such as discarding slow trajectories.
Reward bias: The reward imperfectly represents the behavior we want, such as a verifier that rewards shortcuts.

The focus here is estimator bias, the weighting introduced by the loss, and selection bias from the rollout pipeline.

Before comparing algorithms, it also helps to separate four recurring choices:

Advantage estimation: How do we estimate whether a response or action performed better or worse than expected?
Aggregation: How are token contributions combined and weighted across responses and prompts, including responses of different lengths?
Off-policy correction: How do we account for data generated by an older or different policy?
Admission: Which rollouts reach training at all?

估计优势

With a reward only at the end of a response, the first question is what to compare that reward against. A group baseline uses other responses to the same prompt, whereas a learned critic predicts return from the current prefix. Both aim to reduce gradient variance, but they spend compute differently.

组基线

GRPO (Group Relative Policy Optimization) compares each response’s reward with the mean reward of the full group. It also normalizes rewards and token contributions, introducing bias through the reweighting we’ll examine in the next section. [4]

RLOO (REINFORCE Leave-One-Out) instead excludes each response from its own baseline, averaging only the other responses’ rewards. Under independent on-policy sampling, this preserves the expected gradient. [3]

The cost of a group baseline is generating multiple responses and waiting for their rewards. Larger groups estimate the prompt’s mean reward more precisely, but use compute that could go toward other prompts. For long tasks, an otherwise asynchronous pipeline may still leave completed responses waiting for the slowest member of their group.

学习式评论家

PPO is commonly paired with a learned critic and generalized advantage estimation (GAE). [5]

A critic can provide a baseline at each prefix without waiting for other responses to the same prompt. The cost is an additional value model to fit and run, with extra memory, compute, and tuning. A poorly fitted critic may reduce variance less than a group baseline, or even increase it. Recent work finds that better value baselines can compete with group-relative methods on reasoning tasks, but the benefit depends on critic quality. [6]

Bias depends on how the predictions are used. With on-policy data, subtracting a baseline that depends only on the prefix preserves the expected gradient, even if the prediction is imperfect. Bootstrapping instead replaces part of the sampled return with a value prediction and can introduce bias when the critic is inaccurate. GAE controls this reliance on predictions. With complete episodes, no discounting, and lambda=1, it recovers the sampled return minus the baseline. [2]

The choice is whether better predictions and less group waiting justify the critic’s additional compute and training complexity.

更新加权

Even with the same baseline, normalization can change which prompts and responses drive the update.

GRPO illustrates two such choices. It divides centered rewards by each group’s reward standard deviation, and its original formulation averages token contributions within each response. The first changes relative weights across groups. The second changes relative weights across responses, meaning at the same advantage, each token in a 100-token response gets ten times the coefficient of a token in a 1,000-token response. With a negative advantage, a long incorrect response is penalized less per token than a short one, which can be a driver of growing response length on failures. [4]

Dr.GRPO removes both normalizers, using a fixed scaling constant instead of each response’s length. With fixed group size and the basic on-policy assumptions, its group-centered update is proportional to RLOO’s. This removes those two sources of reweighting without requiring a critic. It does not guarantee lower variance, and clipping or stale data still introduces other approximations. [7]

So far, these comparisons assume fresh, on-policy data. Older rollouts introduce another reason to change their weight in the update.

从不同策略学习

In a working training system, the policy that generated a rollout may differ from the one being updated. This can happen for several reasons:

Policy lag: Training continues while rollouts are generated or queued.
Numerical differences: Quantization and different kernels can produce different probabilities even with the same weights.
Sampling settings: Temperature, top-k, or top-p can change the distribution used to generate tokens.
Retained state: A rollout may continue with new weights while using a KV cache built under older weights.

Some mismatch is a deliberate cost of higher throughput. Async execution reduces waiting for slow trajectories. [1] Retaining KV caches avoids recomputing long prefixes. [8] To correct for the mismatch, we save each token’s probability when it is generated, as recomputing that probability later with the training model may give a different value.

修正不匹配

Off-policy correction compares the behavior policy μ, which generated the response, with the current policy πθ, which is being trained.

In the simple complete-response setting, exact trajectory importance sampling weights a response by:
w(y)=\frac{\pi_\theta(y\mid x)}{\mu(y\mid x)}
=\prod_{t=1}^{T}
\frac{\pi_\theta(y_t\mid x,y_{<t})}
{\mu(y_t\mid x,y_{<t})}

Multiplying the basic gradient estimate by this weight recovers the current-policy gradient in expectation, provided the baseline conditions hold, the behavior probabilities are correct, and the sampler gives positive probability wherever the target does. Top-k or top-p sampling can violate that last condition.

The difficulty is variance, especially for long generations. Multiplying ratios across the response can produce very large or very small weights, leaving a few trajectories dominating the update.

A common approximation is to weight each token’s gradient by that token’s probability ratio, rather than multiplying ratios across the whole response. GRPO uses this token-level form inside its clipped objective. [4] This avoids the long product, but corrects only the action distribution at the recorded prefix. It does not account for how likely the current policy was to reach that prefix or generate the rest of the response that earned the reward.

限制不稳定的更新

Exact correction can be too noisy to use directly. Practical recipes combine approximate correction with limits on policy updates:

Truncated importance sampling caps the correction weight. CISPO, introduced in MiniMax-M1, clips token-level importance weights and treats them as fixed coefficients during differentiation. A token can still contribute a gradient after its weight reaches the cap. This limits extreme weighting at the cost of bias. MiniMax-M1 used an upper cap without a lower floor. [9]

PPO-style clipping removes the incentive to keep changing a token’s probability in an already-favored direction. For example, with a 20% clipping range, a positive-advantage token’s policy-surrogate gradient is zero above a ratio of 1.2, and a negative-advantage token’s is zero below 0.8. Other tokens and loss terms can still move its probability, so this is not a hard constraint on policy movement. [5]

Masking removes selected token or trajectory contributions. IcePop, used in Ring-1T’s reasoning RL, masks tokens whose trainer-to-sampler probability ratio falls outside an allowed interval. This suppresses extreme contributions but discards their signal and generally introduces bias. [10]

GSPO (Group Sequence Policy Optimization) changes the unit of weighting and clipping. It uses the geometric mean of token ratios and clips at the sequence level. Averaging log-ratios rather than summing them moderates the weight fluctuations that long products can produce. The tradeoff is bias relative to the expected-reward gradient: this is no longer exact trajectory correction, and the gradient includes a 1/T factor that changes response weighting even when the policies match. Furthermore, on longer responses, a short stretch of large mismatch can be diluted by the sequence average.

In Qwen’s reported MoE experiments, GSPO improved stability and removed the need for routing replay, which records and reuses expert assignments across policy updates. Avoiding replay saves memory and communication, illustrating why bias alone does not determine training efficiency. [11]

用评分中心化修正漂移

Capping or masking weights controls extreme contributions. Score centering addresses another consequence of mismatch, which is a systematic drift in the expected update. Under mismatch, the trainer’s token score (the gradient of its log probability) can have a nonzero mean under the sampler. Even constant rewards can then produce an update. A small drift nudges the trainer toward the sampler’s distribution, the updated weights are synced back to the sampler, and the bias compounds with every step. Subtracting this mean at each prefix removes that drift, which reward centering across responses does not generally eliminate.

Score Centering Stabilizes Off-policy Reinforcement Learning reports strong results under quantization and gains from combining centering with truncated or masked importance sampling under severe staleness. [12]

The implementation limits logging and computation by storing top-k sampler probabilities and approximating the tail with the rescaled trainer distribution. The correction is then computed over just those k tokens. With k=128, the authors measured less than 1% runtime overhead in their custom implementation. [12]

选择哪些轨迹进入训练

The choices so far determine how a rollout contributes to the update. The pipeline also decides whether to train on it at all. That can save learner time, but it can change the data the model learns from.

When every response in a group receives the same reward, a group-centered advantage is zero. DAPO’s dynamic sampling keeps generating until it fills the batch with groups that have reward variation. This is one component of a broader recipe that also changes clipping, token aggregation, and treatment of overlong responses. Dynamic sampling spends additional generation compute to give the learner a more useful batch. [13]

Filtering changes which prompts reach training, but does not always change the expected gradient’s direction. If groups are sampled independently and averaged equally, and rejected groups contribute exactly zero gradient, removing them changes only the overall scale. This no longer necessarily holds with token-count normalization or another loss, such as a KL penalty, to which the rejected groups still contribute.

Extra sampling is most attractive when zero-advantage groups waste learner capacity. If generation already dominates the cost, the additional rollouts may take more time than they save.

Discarding slow trajectories can introduce selection bias, as those trajectories may carry useful gradients and disproportionately represent harder tasks. A pipeline can appear faster while shifting training toward tasks that finish quickly. Treating a generation timeout as a failed task can also bias the reward, unless that time limit is part of the task’s success criteria.

选择完整方案

The best recipe depends on the tasks, the cost of generating and evaluating rollouts, and the available compute budget. For tasks with terminal, verifiable rewards and affordable groups, a practical starting point combines Dr.GRPO’s normalization choices, a CISPO-style update, and score centering, with asynchronous rollout and training that limits policy lag.

The group baseline can reduce variance without a critic. Dr.GRPO’s normalization choices avoid reweighting by each group’s reward spread or each response’s length. CISPO-style caps control extreme importance weights, and score centering addresses residual drift. Async execution reduces waiting, with regular weight refreshes and limits on queued data.

With that said, we must evaluate the combination on the target workload, as removing specific sources of bias does not guarantee the best learning per GPU-hour.

ScaleRL offers empirical support for evaluating the update and pipeline together. Its reasoning experiments combine asynchronous PipelineRL with CISPO and find that several design choices mainly improve compute efficiency within the full recipe. Its normalization and aggregation differ from those suggested here, and score centering is a later addition to consider. The shared lesson is to judge the complete setup by the learning it delivers for the compute spent. [14]

A few conditions change that starting point:
Little synchronization cost or mismatch: A simpler pipeline may be sufficient. Async execution and extra correction need to justify their overhead.
Substantial policy lag: Refresh weights more often or reduce queueing. Correction cannot make arbitrarily stale data useful, so measure whether the extra throughput still improves learning.
Long or costly rollouts: Group baselines multiply rollout cost per prompt and make finished responses wait on the slowest member. Compare smaller groups, or a learned critic that can train on unfinished trajectories, against the variance reduction the group provides.
Many zero-advantage groups: DAPO-style sampling can improve the learner’s batch, but include the cost of rejected rollouts.

The deciding question is whether training remains stable and reaches the intended quality sooner or with less compute.

阅读原文
📚 相关主题 大语言模型强化学习

订阅 AI Pulse

每天 08:00 · 12:30 · 18:30 · 23:50 更新