DRPO: Smoothing the Path for LLM Reinforcement Learning via Divergence Regularization
Rethinking the Divergence Regularization in LLM RL
The paper introduces Divergence Regularized Policy Optimization (DRPO), a novel reinforcement learning (RL) objective for Large Language Models (LLMs). It replaces the standard ratio-based clipping in PPO/GRPO with a smooth quadratic regularizer based on Binary Total Variation (Binary-TV) to ensure stable off-policy optimization across various architectures and precisions.
TL;DR
Reinforcement Learning (RL) for LLMs is notoriously unstable due to "off-policy" discrepancies between training and inference engines. Divergence Regularized Policy Optimization (DRPO) solves this by ditching the brittle "importance sampling ratio" used in PPO for a smooth regularizer based on absolute probability shifts. This shift provides a bounded, corrective gradient that prevents training collapse—even in extreme low-precision (FP8) environments—and sets new standards for stability in mathematical reasoning tasks.
Problem & Motivation: The "Long-Tail" Ratio Trap
Most mainstream RL algorithms like PPO and GRPO use a ratio-based trust region. The logic is simple: keep the ratio of current policy to behavior policy () near 1.0.
However, the authors point out a fundamental flaw: LLM vocabularies are vast and long-tailed.
- For a rare token (prob: ), a tiny absolute change can cause a massive ratio spike (), triggering unfair gradient clipping.
- For a frequent token (prob: ), a massive shift to results in a modest ratio change (), potentially escaping the trust region unnoticed.
Prior work like DPPO attempted to fix this with a "hard mask" based on probability shift, but it suffered from "gradient death"—once a token crossed the line, it provided no feedback on how to get back.
Methodology: The Core Logic of DRPO
DRPO effectively marries the smooth stability of Simple Policy Optimization (SPO) with the robust geometry of DPPO.
Instead of regularizing the ratio, DRPO applies an advantage-weighted quadratic penalty to the absolute probability shift (). The gradient of the DRPO objective introduces a continuous weight :
Why this works:
- Correction, not just Clipping: When the policy drifts too far (outside ), the gradient weight becomes negative, literally pulling the model back toward the behavior policy.
- Bounded Weights: Unlike SPI/SPO, where weights can explode toward infinity in the "long tail" (low ), DRPO's weights are strictly bounded, ensuring no single token can hijack the entire gradient update.
Graph: Note how DRPO's weight (bottom right) remains bounded across all probabilities, unlike SPO (top right) which spikes at low probabilities.
Experiments & Results: Stability Meets SOTA
The team tested DRPO across various scales (1.5B to 35B parameters) and precision settings (BF16, FP8).
Key Findings:
- FP8 Resilience: In "End-to-End FP8" training—a setting that usually causes RL to collapse—DRPO remained perfectly stable while GRPO and SPO failed.
- Superior Convergence: On the AIME (American Invitational Mathematics Examination) benchmarks, DRPO reached higher peak accuracy faster than its predecessors.
Figure: DRPO consistently maintains a superior accuracy curve across various model backbones compared to hard-clipped (GRPO) or unregularized methods.
Deep Insights: The "Gradient-First" View
A crucial takeaway from this research is the Ablation on Advantage Weighting. Many researchers use a pure KL penalty, but the authors found that weighting the penalty by the absolute advantage () is vital. This ensures the trust-region boundary stays fixed regardless of whether the reward is +1 or +100.
Ablation Study: Removing the advantage weight (orange) leads to significant performance degradation compared to the full DRPO (blue).
Conclusion
DRPO represents a shift in thinking from "Which divergence should I minimize?" to "What does the gradient weight look like?" By choosing an -style penalty on probability space, DRPO provides a robust, smooth, and corrective framework for LLM RL. It is particularly valuable for groups training at scale with low precision (FP8/FP4), where numerical noise makes traditional ratio-clipping too brittle.
Limitations: While the Binary-TV approximation is efficient, it only looks at the sampled token. Future work might explore "Top-K TV" for even more precise distributional control.
