Token Bill

tool-churn.jsonl · 1 run · 14 calls · models: claude-sonnet-5 · {{REPORT_DATE}}

~67%≈ of billed input tokens went to re-sending bytes the model had already seen; the fix below recovers an estimated $0.0882 of $0.12.

Run demo-tool-churn-seed7

14 calls · cache read 0 · cache write 0 · uncached input 56,734 · output 993 tokens · billed $0.12 · redundant input 66.8%≈

cache readcache writeuncached inputoutput01.6k3.1k4.7k6.2k024681012call indexbilled tokens
Billed tokens per call (exact, from the trace's usage fields): cache reads and writes, uncached input, output.
as-billed$0.12no-cache$0.12optimal-cache$0.0641fixed-cache$0.0352
Dollars under each scenario. as-billed and no-cache are exact arithmetic on billed usage; optimal-cache and fixed-cache are simulations on the approx char basis≈.

Cache breakers (1)

tool-churn first at call index 4 · recovers ≈ $0.0882

Fix: send tool definitions in one fixed order on every call (sort them once at startup); reordering rewrites the cached prefix

tool order changed at call 4: first seen ['read_file', 'edit_file', 'run_command', 'search_code'] -> ['edit_file', 'run_command', 'search_code', 'read_file']

Methodology

  1. Exact vs approximate. Every dollar figure and token total comes from the trace's real billed usage fields (or exact arithmetic on them). Numbers marked ≈ rest on char-based attribution (len/3.7), scaled so segments sum to each call's billed total — useful for proportions, never presented as billed.
  2. Cache simulation rules. 300-second cache TTL, refreshed on read (sliding-window assumption, documented); cache entries are per-model; writes billed at the model's write premium only when a later call in the replay actually reads the entry (an optimal policy never caches what nothing reads back), so optimal-cache never exceeds no-cache; a prefix must meet the model's minimum cacheable length; the simulator places a single cache breakpoint at the end of messages each call (optimal placement). It models the provider's documented rules, not undocumented server behavior.
  3. Redundancy. For each call after the first: the byte-identical rendered prefix shared with the previous call, valued at the call's billed input scaled by char fraction; input already served as cache reads is subtracted (cache reads are cheap — they are not waste).
  4. Scenarios. as-billed = ground truth from usage; no-cache = all input at the full uncached rate; optimal-cache = replay under the documented cache rules; fixed-cache = optimal-cache after neutralizing the detected breakers.
  5. Pricing. Bundled table shipped with tokenbill {{TOKENBILL_VERSION}}, verified 2026-09 against the provider price list. Cache reads are 0.10× base input (0.025× on claude-fable-5-1); dated snapshot ids are priced as their base model. Re-verify before release-grade accounting.