Find out when your Claude prompt cache stops working on Bedrock.
Cached prompts cost about a tenth as much to read. When caching breaks, requests still succeed and nothing errors. Usually the bill is the first sign. CacheCanary checks your requests in CI, tells you why a call missed the cache, and shows hit rates from your Bedrock logs.
Free and open source (Apache 2.0). Runs on your machine. Nothing is sent anywhere.
What it looks like
Save a request your app sends to Bedrock as JSON and point CacheCanary at it.
$ cachecanary lint request.json [warn] dynamic-in-prefix: Found ISO date text inside the cached prefix. If it changes per request, every call misses. (system[0]) $ cachecanary diff yesterday.json today.json system-changed: The system prompt changed inside the cached prefix (look for dates, IDs or per-user text). (system[0]) $ cachecanary lint request.json --model us.anthropic.claude-haiku-4-5-20251001-v1:0 [error] prefix-too-short: Prefix up to this checkpoint is ~1892 tokens; claude-haiku-4-5 needs at least 4096. The request succeeds but nothing is cached. (system[0])
The same prompt caches fine on Sonnet 4.6, which needs 1,024 tokens, and never caches on Haiku 4.5, which needs 4,096.
Why caching breaks
Often nobody changed the prompt on purpose. These are the usual causes:
- A library upgrade moved the system prompt or dropped the cache marker. This happened with LiteLLM in July 2026: hit rates fell from about 90% to 25-45% and spend went up 2-3x for six days.
- You moved to a new model ID or an inference profile ARN, and your library no longer knows it can cache.
- A small edit put today's date, a request ID or the user's name into the system prompt.
- Tools are built in a different order on each request.
- An agent made a dozen parallel tool calls, and Bedrock only looks back about 20 blocks for the cached part.
Anthropic added cache diagnostics to its own API, but it doesn't cover Bedrock. That's the gap CacheCanary fills.
Put it in CI
Problems show up as notes on the pull request, next to the file that caused them.
- uses: Haarris/cachecanary@v0
with:
command: lint
args: tests/fixtures/agent_request.json --model us.anthropic.claude-sonnet-4-6
# Or call Bedrock for real (needs AWS credentials)
- uses: Haarris/cachecanary@v0
with:
command: probe
args: tests/fixtures/agent_request.json --model us.anthropic.claude-sonnet-4-6
Tested on real Bedrock
The main cases were run against Amazon Bedrock in October 2026 with made-up prompts: cache hits on both APIs, both TTLs, streaming, model minimums, a date in the system prompt, reordered tools, a model switch and the 20-block lookback.
Coming next
A hosted dashboard: cache hit rate and wasted spend per app, an alert when the rate drops, and the reason for each miss. It runs inside your own AWS account, so prompts never leave it.
If you'd use it, email hello@cachecanary.com. I read every message.