Trackmania reward decomposition and terminal-safe shaping
Ordinary reward terms remain separate from two potential-based shaping terms; terminal transitions use an absorbing zero potential.
Validated signals
Ordinary and terminal reward terms
Potential-based shaping
Transition result
Telemetry delta
Δt · path progress ·
velocity · steering
Track references
lap length · tangent ·
human pace
Terminal contract
finish · off-track ·
time or stall limit
Reward configuration
finite scales · shared
reward_gamma
Time cost
−time_scale · Δt
Direct
progress
lap_scale · Δd
/ L
Projected
velocity
scale · v∥ · Δt
Speed and
steering
speed² bonus
steer Δ penalty
Collision
cooldown-gated
penalty
Terminal base
finish bonus or
failure penalty
Time-attack
finish term
finish timing
bonus
Ordinary subtotal
per-step + terminal
terms
Progress potential
Φp = wp · d / L
Pace potential
Φpace(s) = −wd · clipped
time debt
Non-terminal shaping
F = γ Φ(s′) − Φ(s), for
both potentials
Terminal reset
Φ(terminal) = 0; emit
−Φ(previous) for both
potentials
Total reward
ordinary + progress
PBRS + pace PBRS
RewardResult
breakdown
named terms · potential
· pace debt · collision
Learner transition
scalar reward ·
terminated · reason
Runtime invariant: reward_gamma must equal training.gamma so the implemented PBRS difference uses the learner's discount.
Direct progress is an intentional ordinary term. Progress PBRS is reported separately, and terminal/time-attack diagnostics do not double-count the
base finish reward.
Download editable diagram