~89% of billed input tokens went to re-sending bytes the model had already seen; the fix below recovers an estimated $0.0882 of $0.12.
bundled demo scenario 'no-cache' (seed 7) | 1 run | 14 calls | models: claude-sonnet-5
[synthetic demo data: bundled scenarios with planted waste]

Run demo-no-cache-seed7
  billed tokens    cache read 0 | cache write 0 | uncached input 56,734 | output 993
  billed dollars   $0.12  (cache read $0.00 | cache write $0.00 | uncached input $0.11 | output $0.0099)
  redundant input  ~89.3% of billed input tokens re-sent (approx)
  scenarios
    as-billed        $0.12  ########################
    no-cache         $0.12  ########################
    optimal-cache  $0.0352  #######
    fixed-cache    $0.0352  #######
    note (as-billed): exact: real billed usage priced at published rates (ground truth)
    note (no-cache): counterfactual: every billed input token repriced at the full uncached rate (no cache reads, no write premium)
    note (optimal-cache): simulated (approx): documented cache rules — 300s TTL sliding on read, min-cacheable gate, one breakpoint at end of messages; char-based token split scaled to billed totals
    note (fixed-cache): simulated (approx): optimal-cache rules over the breaker-repaired rendering; billed usage totals reused for the token split
  breakers
    missing-breakpoint | first at call index 1 | recovers ~$0.0882
      fix: add a cache_control breakpoint (for example on the last message); the stable prefix already meets the minimum cacheable length
      evidence: calls 0->1 share a byte-stable prefix of ~1641 approx tokens (min cacheable 1024) but cache_breakpoints=0 and billed cac... [truncated, 136 chars total]

approx (~): char-based attribution scaled to billed totals; dollar and token totals come from real billed usage.
next: tokenbill demo -o report.html writes the full HTML report; then record a real agent — see "Your first real trace" in the README
