Trackmania reward decomposition and terminal-safe shaping Ordinary reward terms remain separate from two potential-based shaping terms; terminal transitions use an absorbing zero potential. Validated signals Ordinary and terminal reward terms Potential-based shaping Transition result Telemetry delta Δt · path progress ·velocity · steering Track references lap length · tangent ·human pace Terminal contract finish · off-track ·time or stall limit Reward configuration finite scales · sharedreward_gamma Time cost −time_scale · Δt Directprogress lap_scale · Δd/ L Projectedvelocity scale · v∥ · Δt Speed andsteering speed² bonussteer Δ penalty Collision cooldown-gatedpenalty Terminal base finish bonus orfailure penalty Time-attackfinish term finish timingbonus Ordinary subtotal per-step + terminalterms Progress potential Φp = wp · d / L Pace potential Φpace(s) = −wd · clippedtime debt Non-terminal shaping F = γ Φ(s′) − Φ(s), forboth potentials Terminal reset Φ(terminal) = 0; emit−Φ(previous) for bothpotentials Total reward ordinary + progressPBRS + pace PBRS RewardResultbreakdown named terms · potential· pace debt · collision Learner transition scalar reward ·terminated · reason Runtime invariant: reward_gamma must equal training.gamma so the implemented PBRS difference uses the learner's discount. Direct progress is an intentional ordinary term. Progress PBRS is reported separately, and terminal/time-attack diagnostics do not double-count thebase finish reward.