Token Bill

timestamp.jsonl · 1 run · 14 calls · models: claude-sonnet-5 · {{REPORT_DATE}}

~18%≈ of billed input tokens went to re-sending bytes the model had already seen; the fix below recovers an estimated $0.0884 of $0.12.

Run demo-timestamp-seed7

14 calls · cache read 0 · cache write 0 · uncached input 56,880 · output 993 tokens · billed $0.12 · redundant input 17.7%≈

cache readcache writeuncached inputoutput01.6k3.1k4.7k6.2k024681012call indexbilled tokens
Billed tokens per call (exact, from the trace's usage fields): cache reads and writes, uncached input, output.
as-billed$0.12no-cache$0.12optimal-cache$0.12fixed-cache$0.0353
Dollars under each scenario. as-billed and no-cache are exact arithmetic on billed usage; optimal-cache and fixed-cache are simulations on the approx char basis≈.

Cache breakers (1)

volatile-system first at call index 1 · recovers ≈ $0.0884

Fix: move the volatile value (timestamp/UUID/counter) out of the system prompt — inject it in the latest user message instead

system chars [355:374] at call 1: '...reen.\nSession: [session 2026-07-26 14:03:00]\n\nRepository layout:\n  ...' -> '...reen.\nSession: [session 2026-07-26 14:03:01]\n\nRepository layout:\n  ...'

Methodology

  1. Exact vs approximate. Every dollar figure and token total comes from the trace's real billed usage fields (or exact arithmetic on them). Numbers marked ≈ rest on char-based attribution (len/3.7), scaled so segments sum to each call's billed total — useful for proportions, never presented as billed.
  2. Cache simulation rules. 300-second cache TTL, refreshed on read (sliding-window assumption, documented); cache entries are per-model; writes billed at the model's write premium only when a later call in the replay actually reads the entry (an optimal policy never caches what nothing reads back), so optimal-cache never exceeds no-cache; a prefix must meet the model's minimum cacheable length; the simulator places a single cache breakpoint at the end of messages each call (optimal placement). It models the provider's documented rules, not undocumented server behavior.
  3. Redundancy. For each call after the first: the byte-identical rendered prefix shared with the previous call, valued at the call's billed input scaled by char fraction; input already served as cache reads is subtracted (cache reads are cheap — they are not waste).
  4. Scenarios. as-billed = ground truth from usage; no-cache = all input at the full uncached rate; optimal-cache = replay under the documented cache rules; fixed-cache = optimal-cache after neutralizing the detected breakers.
  5. Pricing. Bundled table shipped with tokenbill {{TOKENBILL_VERSION}}, verified 2026-09 against the provider price list. Cache reads are 0.10× base input (0.025× on claude-fable-5-1); dated snapshot ids are priced as their base model. Re-verify before release-grade accounting.