~75% of this run's rollout compute produced no learning signal = 19 GPU-hours = about $57 at $2.99/GPU-hour
Estimated from logged step time, assuming 70% of step wall-clock is rollout generation. Override the rate with --gpu-hour-cost.
Run at a glance
training reward
1.05
held-out score
0.9206
entropy
0.1482
zero-variance groups
0.6793
KL to reference
0.3653
response length
506
clip fraction
0.1025
grad norm
0.2645
Findings
CRITZero-variance groups (wasted rollouts)
The task is too easy: most groups are all-correct and contribute no gradient. 75% of your rollout budget produces exactly zero policy gradient.
Raise task difficulty -- filter out prompts the current policy already solves with pass rate > 0.9 and refresh the training pool.
Adopt a curriculum keyed on measured pass rate; target the 0.3-0.7 band where group variance, and therefore gradient, is maximal.
Enable dynamic sampling (DAPO): keep resampling prompts until each group has non-zero reward variance. verl: `algorithm.filter_groups.enable=True`; TRL: see the dynamic-sampling issue tracker for the current flag.
Evidence
measured zero-variance group fraction over the last quarter of the run: 75.2% (median of 100 steps)
effective sample efficiency: 24.8% -- you pay for 4.0 rollouts per rollout that produces gradient
trend is rising (+2.59e-03/step, Mann-Kendall p<1e-16), so the waste is getting worse
mean pass rate is 95% -- the model has outgrown this data
References
Yu et al., DAPO: An Open-Source LLM Reinforcement Learning System at Scale (arXiv:2503.14476) -- dynamic sampling
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation (arXiv:2605.21125)
TRL logs this directly as `frac_reward_zero_std`
WARNEntropy collapse (exploration exhausted)
Entropy is collapsing and will hit the exploration floor in roughly 226 steps. Reward has already plateaued.
Investigate the abrupt drop first. Diff your config around that step and check for a resumed checkpoint or a changed reward function before tuning anything.
Decouple the clip bounds (DAPO 'clip-higher'): raise `clip_eps_high` to ~0.28 while leaving `clip_eps_low` at 0.2. Low-probability tokens are currently clipped away before they can be reinforced, which is a primary driver of entropy collapse.
Add a small entropy bonus (start at 1e-3 and tune by an order of magnitude). Treat it as a floor-holding device, not a cure -- a large bonus trades collapse for a different instability.
Reward has plateaued while entropy keeps falling: this run is finished learning. Stop it, keep the best checkpoint, and spend the remaining budget on harder data.
Raise rollout temperature slightly, or sample with a higher top-p, to restore group diversity without touching the loss.
Evidence
entropy fell from 0.6720 to 0.1529 over 399 steps (77% of the initial exploration budget spent)
still falling in the recent half at -5.999e-04/step (Mann-Kendall p<1e-16, tau=-0.88)
extrapolating the exponential fit, entropy reaches the floor (0.0672) in about 226 steps -- roughly 2.6 GPU-wall-clock hours from now
abrupt drop detected at step 78: 0.5824 -> 0.2323 (17.0 robust sigma). A cliff rather than a slope usually means a config change, a resume from checkpoint, or a reward-function bug landing mid-run.
held-out eval trend over the same window: +8.130e-05/step (p=0.536) -- reward has stopped improving, so this decay is buying nothing
References
Cui et al., The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models (Clip-Cov / KL-Cov)
Yu et al., DAPO (arXiv:2503.14476) -- clip-higher as an entropy remedy
Understanding and Preventing Entropy Collapse in RLVR with On-Policy Regularisation (arXiv:2605.11491)
INFOPPO clip saturation / asymmetry
Clip fraction is 10%, on the high side.
Evidence
10.1% of token updates are being clipped away (trend +1.00e-04/step, p<1e-16)
Schulman et al., Proximal Policy Optimization Algorithms (arXiv:1707.06347)
Yu et al., DAPO (arXiv:2503.14476) -- decoupled clip bounds ('clip-higher')
INFOReward-eval divergence (verifier gaming)
Training reward and held-out score have drifted apart (gap +2.58, rho=+0.96). Worth an eyeball before you trust the reward curve.
Read 20 high-reward rollouts end to end. Reward hacks are almost always obvious on inspection and almost never visible in aggregate metrics.
Audit the verifier, not the model. For code tasks, check whether your tests accept known-wrong patches; for rubric rewards, check for keyword stuffing.
Hold out a second verifier that grades the same task differently (isomorphic verification). A gap that exists against one grader and not another localises the exploit to the grader.
Add an explicit anti-hack penalty for the specific exploit once you have identified it, and re-measure the gap rather than assuming it closed.
Roll back to the checkpoint before the gap opened -- later checkpoints have already been shaped by the exploit.
Evidence
normalised training reward has moved +10.01 from its start; held-out score has moved +7.57 -- hacking gap +2.58
rank correlation between training reward and held-out score: rho=+0.96 over 17 evaluations
the gap is widening (+7.59e-03/step, p=0.000463)
References
Auditing Reward Hackability in Code RL Training Environments (arXiv:2606.16062)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking (arXiv:2604.15149)
Reward Hacking in the Era of Large Models: Mechanisms and Emergent Behaviours (arXiv:2604.13602)