Token Bill

bundled demo scenarios (seed 11) · 4 runs · 56 calls · models: claude-sonnet-5 · {{REPORT_DATE}}

~42%≈ of billed input tokens went to re-sending bytes the model had already seen; the three fixes below recover an estimated $0.35 of $0.52.

Run demo-well-behaved-seed11

14 calls · cache read 65,974 · cache write 7,955 · uncached input 0 · output 1,155 tokens · billed $0.0446 · redundant input 0.0%≈

cache readcache writeuncached inputoutput02.0k4.0k6.0k8.1k024681012call indexbilled tokens
Billed tokens per call (exact, from the trace's usage fields): cache reads and writes, uncached input, output.
as-billed$0.0446no-cache$0.16optimal-cache$0.0444fixed-cache$0.0444
Dollars under each scenario. as-billed and no-cache are exact arithmetic on billed usage; optimal-cache and fixed-cache are simulations on the approx char basis≈.

Cache breakers

None detected.

Run demo-timestamp-seed11

14 calls · cache read 0 · cache write 0 · uncached input 74,080 · output 1,155 tokens · billed $0.16 · redundant input 13.6%≈

cache readcache writeuncached inputoutput02.0k4.0k6.1k8.1k024681012call indexbilled tokens
Billed tokens per call (exact, from the trace's usage fields): cache reads and writes, uncached input, output.
as-billed$0.16no-cache$0.16optimal-cache$0.16fixed-cache$0.0445
Dollars under each scenario. as-billed and no-cache are exact arithmetic on billed usage; optimal-cache and fixed-cache are simulations on the approx char basis≈.

Cache breakers (1)

volatile-system first at call index 1 · recovers ≈ $0.12

Fix: move the volatile value (timestamp/UUID/counter) out of the system prompt — inject it in the latest user message instead

system chars [355:374] at call 1: '...reen.\nSession: [session 2026-07-26 14:03:00]\n\nRepository layout:\n  ...' -> '...reen.\nSession: [session 2026-07-26 14:03:01]\n\nRepository layout:\n  ...'

Run demo-tool-churn-seed11

14 calls · cache read 0 · cache write 0 · uncached input 73,929 · output 1,155 tokens · billed $0.16 · redundant input 66.3%≈

cache readcache writeuncached inputoutput02.0k4.0k6.0k8.1k024681012call indexbilled tokens
Billed tokens per call (exact, from the trace's usage fields): cache reads and writes, uncached input, output.
as-billed$0.16no-cache$0.16optimal-cache$0.0826fixed-cache$0.0444
Dollars under each scenario. as-billed and no-cache are exact arithmetic on billed usage; optimal-cache and fixed-cache are simulations on the approx char basis≈.

Cache breakers (1)

tool-churn first at call index 4 · recovers ≈ $0.11

Fix: send tool definitions in one fixed order on every call (sort them once at startup); reordering rewrites the cached prefix

tool order changed at call 4: first seen ['read_file', 'edit_file', 'run_command', 'search_code'] -> ['edit_file', 'run_command', 'search_code', 'read_file']

Run demo-no-cache-seed11

14 calls · cache read 0 · cache write 0 · uncached input 73,929 · output 1,155 tokens · billed $0.16 · redundant input 89.2%≈

cache readcache writeuncached inputoutput02.0k4.0k6.0k8.1k024681012call indexbilled tokens
Billed tokens per call (exact, from the trace's usage fields): cache reads and writes, uncached input, output.
as-billed$0.16no-cache$0.16optimal-cache$0.0444fixed-cache$0.0444
Dollars under each scenario. as-billed and no-cache are exact arithmetic on billed usage; optimal-cache and fixed-cache are simulations on the approx char basis≈.

Cache breakers (1)

missing-breakpoint first at call index 1 · recovers ≈ $0.11

Fix: add a cache_control breakpoint (for example on the last message); the stable prefix already meets the minimum cacheable length

calls 0->1 share a byte-stable prefix of ~1641 approx tokens (min cacheable 1024) but cache_breakpoints=0 and billed cache activity is 0

Methodology

  1. Exact vs approximate. Every dollar figure and token total comes from the trace's real billed usage fields (or exact arithmetic on them). Numbers marked ≈ rest on char-based attribution (len/3.7), scaled so segments sum to each call's billed total — useful for proportions, never presented as billed.
  2. Cache simulation rules. 300-second cache TTL, refreshed on read (sliding-window assumption, documented); cache entries are per-model; writes billed at the model's write premium only when a later call in the replay actually reads the entry (an optimal policy never caches what nothing reads back), so optimal-cache never exceeds no-cache; a prefix must meet the model's minimum cacheable length; the simulator places a single cache breakpoint at the end of messages each call (optimal placement). It models the provider's documented rules, not undocumented server behavior.
  3. Redundancy. For each call after the first: the byte-identical rendered prefix shared with the previous call, valued at the call's billed input scaled by char fraction; input already served as cache reads is subtracted (cache reads are cheap — they are not waste).
  4. Scenarios. as-billed = ground truth from usage; no-cache = all input at the full uncached rate; optimal-cache = replay under the documented cache rules; fixed-cache = optimal-cache after neutralizing the detected breakers.
  5. Pricing. Bundled table shipped with tokenbill {{TOKENBILL_VERSION}}, verified 2026-09 against the provider price list. Cache reads are 0.10× base input (0.025× on claude-fable-5-1); dated snapshot ids are priced as their base model. Re-verify before release-grade accounting.