rldoctor 0.1.0 — sim:saturated_groups

400 steps · 15 metrics · GRPO · Qwen2.5-7B-Instruct · group size 8 · 8× H100 · source: simulated

Run at a glance

training reward
1.05
held-out score
0.9206
entropy
0.1482
zero-variance groups
0.6793
KL to reference
0.3653
response length
506
clip fraction
0.1025
grad norm
0.2645

Findings

CRITZero-variance groups (wasted rollouts)
The task is too easy: most groups are all-correct and contribute no gradient. 75% of your rollout budget produces exactly zero policy gradient.
  1. Raise task difficulty -- filter out prompts the current policy already solves with pass rate > 0.9 and refresh the training pool.
  2. Adopt a curriculum keyed on measured pass rate; target the 0.3-0.7 band where group variance, and therefore gradient, is maximal.
  3. Enable dynamic sampling (DAPO): keep resampling prompts until each group has non-zero reward variance. verl: `algorithm.filter_groups.enable=True`; TRL: see the dynamic-sampling issue tracker for the current flag.
Evidence
  • measured zero-variance group fraction over the last quarter of the run: 75.2% (median of 100 steps)
  • effective sample efficiency: 24.8% -- you pay for 4.0 rollouts per rollout that produces gradient
  • trend is rising (+2.59e-03/step, Mann-Kendall p<1e-16), so the waste is getting worse
  • mean pass rate is 95% -- the model has outgrown this data
References
  • Yu et al., DAPO: An Open-Source LLM Reinforcement Learning System at Scale (arXiv:2503.14476) -- dynamic sampling
  • Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation (arXiv:2605.21125)
  • TRL logs this directly as `frac_reward_zero_std`
WARNEntropy collapse (exploration exhausted)
Entropy is collapsing and will hit the exploration floor in roughly 226 steps. Reward has already plateaued.
  1. Investigate the abrupt drop first. Diff your config around that step and check for a resumed checkpoint or a changed reward function before tuning anything.
  2. Decouple the clip bounds (DAPO 'clip-higher'): raise `clip_eps_high` to ~0.28 while leaving `clip_eps_low` at 0.2. Low-probability tokens are currently clipped away before they can be reinforced, which is a primary driver of entropy collapse.
  3. Add a small entropy bonus (start at 1e-3 and tune by an order of magnitude). Treat it as a floor-holding device, not a cure -- a large bonus trades collapse for a different instability.
  4. Reward has plateaued while entropy keeps falling: this run is finished learning. Stop it, keep the best checkpoint, and spend the remaining budget on harder data.
  5. Raise rollout temperature slightly, or sample with a higher top-p, to restore group diversity without touching the loss.
Evidence
  • entropy fell from 0.6720 to 0.1529 over 399 steps (77% of the initial exploration budget spent)
  • still falling in the recent half at -5.999e-04/step (Mann-Kendall p<1e-16, tau=-0.88)
  • extrapolating the exponential fit, entropy reaches the floor (0.0672) in about 226 steps -- roughly 2.6 GPU-wall-clock hours from now
  • abrupt drop detected at step 78: 0.5824 -> 0.2323 (17.0 robust sigma). A cliff rather than a slope usually means a config change, a resume from checkpoint, or a reward-function bug landing mid-run.
  • held-out eval trend over the same window: +8.130e-05/step (p=0.536) -- reward has stopped improving, so this decay is buying nothing
References
  • Cui et al., The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models (Clip-Cov / KL-Cov)
  • Yu et al., DAPO (arXiv:2503.14476) -- clip-higher as an entropy remedy
  • Understanding and Preventing Entropy Collapse in RLVR with On-Policy Regularisation (arXiv:2605.11491)
INFOPPO clip saturation / asymmetry
Clip fraction is 10%, on the high side.
Evidence
  • 10.1% of token updates are being clipped away (trend +1.00e-04/step, p<1e-16)
  • one-sided clipping: lower bound 5.16%, upper bound 4.72% (ratio 1.1x)
References
  • Schulman et al., Proximal Policy Optimization Algorithms (arXiv:1707.06347)
  • Yu et al., DAPO (arXiv:2503.14476) -- decoupled clip bounds ('clip-higher')
INFOReward-eval divergence (verifier gaming)
Training reward and held-out score have drifted apart (gap +2.58, rho=+0.96). Worth an eyeball before you trust the reward curve.
  1. Read 20 high-reward rollouts end to end. Reward hacks are almost always obvious on inspection and almost never visible in aggregate metrics.
  2. Audit the verifier, not the model. For code tasks, check whether your tests accept known-wrong patches; for rubric rewards, check for keyword stuffing.
  3. Hold out a second verifier that grades the same task differently (isomorphic verification). A gap that exists against one grader and not another localises the exploit to the grader.
  4. Add an explicit anti-hack penalty for the specific exploit once you have identified it, and re-measure the gap rather than assuming it closed.
  5. Roll back to the checkpoint before the gap opened -- later checkpoints have already been shaped by the exploit.
Evidence
  • normalised training reward has moved +10.01 from its start; held-out score has moved +7.57 -- hacking gap +2.58
  • rank correlation between training reward and held-out score: rho=+0.96 over 17 evaluations
  • the gap is widening (+7.59e-03/step, p=0.000463)
References
  • Auditing Reward Hackability in Code RL Training Environments (arXiv:2606.16062)
  • LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking (arXiv:2604.15149)
  • Reward Hacking in the Era of Large Models: Mechanisms and Emergent Behaviours (arXiv:2604.13602)

Checks

✓ length_pathology✓ kl_drift✓ gradient_pathology✓ reward_composition✓ plateau