{% extends "base.html" %}
{# CUI // SP-CTI — Cache Savings Dashboard (D-CACHE-VIS-1) #}
{# NIST 800-53: SC-28 (Protection at Rest), AU-12 (Audit Record Generation), SA-11 (Developer Testing) #}
{% block title %}Cache Savings — ICDEV™{% endblock %}
{% set iqe_canvas = "cache_savings" %}
{% set iqe_api_route = "/api/core/iqe-query" %}
{% set iqe_title = "Query Cache Stats" %}
{% set iqe_examples = [
{"label": "Hit rate by function", "query": "SELECT function, hit_rate_pct FROM cache.stats ORDER BY hit_rate_pct DESC"},
{"label": "Top functions by hits", "query": "SELECT function, total_hits, avoided_calls FROM cache.stats ORDER BY total_hits DESC LIMIT 10"},
{"label": "Cost savings breakdown", "query": "SELECT function, cost_usd_saved FROM cache.stats ORDER BY cost_usd_saved DESC"},
{"label": "Prefix cache by provider", "query": "SELECT provider, status, cached_share_pct, usd_saved FROM cache.by_provider ORDER BY calls DESC"},
{"label": "Providers with no traffic", "query": "SELECT provider, capability FROM cache.by_provider WHERE status = 'no_data'"},
{"label": "Spend that did not ship", "query": "SELECT task_id, verdict, status, cost_usd FROM cache.spend WHERE verdict != 'shipped'"},
{"label": "Cards with unpriced dispatches", "query": "SELECT task_id, verdict, dispatches, unpriced_dispatches FROM cache.spend WHERE unpriced_dispatches > 0"},
] %}
{% block content %}
Cache Savings
{{ "ENABLED" if stats.enabled else "DISABLED" }}
backend: {{ stats.backend }}
JSON API ↗
{# None means the cache was never ASKED in the window — there is no rate.
Rendering 0% there says the cache failed every request it was given,
which is the opposite of what happened (cch-obs-03). #}
{% if stats.summary.hit_rate_pct is none %}
—
Cache Hit Rate
not measured — no requests in window
{% else %}
{{ stats.summary.hit_rate_pct|round(1) }}%
Cache Hit Rate
{% endif %}
{{ "{:,}".format(stats.summary.total_entries) }}
Cached Entries
{{ "{:,}".format(stats.summary.total_hits) }} total hits
{{ "{:,}".format(stats.summary.cache_read_tokens) }}
Context Cache Reads
{{ "{:,}".format(stats.summary.cache_write_tokens) }} written
${{ "%.4f"|format(stats.summary.total_usd_saved) }}
Estimated Savings (USD)
resp: ${{ "%.4f"|format(stats.summary.resp_cache_usd_saved) }} +
ctx: ${{ "%.4f"|format(stats.summary.context_cache_usd_saved) }}
{% if stats.by_function %}
By Function
| Function |
Entries |
Hit Rate |
Avoided Calls |
Read Tokens |
Cost Saved |
{% for fn in stats.by_function %}
| {{ fn.function }} |
{{ "{:,}".format(fn.total_entries) }} |
{% if fn.hit_rate_pct is none %}
—
{% else %}
{{ fn.hit_rate_pct }}%
{% endif %}
|
{{ "{:,}".format(fn.avoided_calls) }} |
{{ "{:,}".format(fn.cache_read_tokens) }} |
${{ "%.4f"|format(fn.cost_usd_saved) }} |
{% endfor %}
{% else %}
No cached entries yet. Cache activates after the first LLM invocation on an enabled function.
{% endif %}
{# ------------------------------------------------------------------
Per-provider PREFIX cache effectiveness (cch-obs-01).
Everything above this point describes the RESPONSE cache — calls avoided
outright — which is legitimately one number. This section describes
PREFIX caching, which is not: the providers do different things, so one
rate over the mix is a blur rather than a summary.
The colour rules are the whole point. Only `caching` and `no_cache_hits`
are on a red/amber/green scale, because only they measured something.
`no data` and `not reported` are neutral and say why — a provider with no
traffic and a provider that reports no counters are not failing caches.
------------------------------------------------------------------ #}
Prefix Cache by Provider
{% if by_provider and by_provider.measurable %}
last {{ by_provider.window_days }}d · {{ by_provider.window_start }} to {{ by_provider.window_end }}
{% endif %}
JSON API ↗
{% if not by_provider or not by_provider.measurable %}
Unmeasurable.
{{ by_provider.unmeasurable_reason if by_provider else "the per-provider view did not load." }}
This is deliberately not shown as 0% on every provider — a database with no
operating history and a platform where caching is broken would look identical.
{% else %}
| Provider |
Caching |
Status |
Calls |
Cached In |
Uncached In |
Cached Share |
Trend |
Saved |
{% for p in by_provider.providers %}
{% set measured = p.status in ['caching', 'no_cache_hits'] %}
| {{ p.provider }} |
{{ p.capability|replace('_', ' ') }} |
{{ p.status_label }}
|
{{ "{:,}".format(p.calls) }} |
{{ "{:,}".format(p.cached_input_tokens) }} |
{{ "{:,}".format(p.uncached_input_tokens) }} |
{%- if p.cached_share_pct is none %}
{# NOT 0%. There is no rate to report here, and printing one would
be the exact conflation this card was opened to remove. #}
—
{%- else %}
{{ p.cached_share_pct }}%
{%- endif %}
|
{%- if p.trend.direction == 'improved' %}
▲ {{ p.trend.delta_pct_points }} pp
{%- elif p.trend.direction == 'worsened' %}
▼ {{ p.trend.delta_pct_points }} pp
{%- elif p.trend.direction == 'flat' %}
flat
{%- else %}
no baseline
{%- endif %}
|
{%- if p.usd_saved is none %}
{# A local provider has no bill, so it has no dollars to save.
$0.00 here would read as "caching failed" for a cache that is
working fine and simply is not billed. Latency instead. #}
{# A local provider's latency is its only observable effect —
but only when it was actually called. "0 ms avg" for a
provider with no traffic is the same fabricated zero this
table exists to remove, one column over. #}
{%- if p.usd_basis == 'local' and p.calls > 0 %}{{ p.avg_latency_ms|round(0)|int }} ms avg
{%- elif p.usd_basis == 'local' %}—
{%- else %}not priced{% endif -%}
{%- else %}
${{ "%.4f"|format(p.usd_saved) }}
{%- endif %}
|
{% endfor %}
{{ by_provider.totals.providers_caching }} caching ·
{{ by_provider.totals.providers_no_cache_hits }} measured zero ·
{{ by_provider.totals.providers_unreported }} not reported ·
{{ by_provider.totals.providers_no_data }} no data.
Total saved ${{ "%.4f"|format(by_provider.totals.usd_saved_total or 0) }}
— {{ by_provider.totals.usd_saved_basis }}.
There is deliberately no combined hit rate on this table. Providers disagree on
whether reported input tokens already include the cached ones (Anthropic and Bedrock
report them separately; OpenAI and Azure report cached tokens as a subset), so one
average over the mix would double-count every cached token from half of them.
Claims live in args/cache_effectiveness.yaml.
{% endif %}
{# ------------------------------------------------------------------
SPEND BY CARD (xrv-cost-04)
A different question from everything above it. The tables above ask what
CACHING saved; this asks what the board SPENT and whether that spend
shipped. They are never added together and never share a rate.
Every empty state is rendered IN WORDS. $0.00 on a cost panel reads as
"this was free", which is the one thing an unattributed board must never
be mistaken for — the cards may well have cost money, nothing attributed
it. `unpriced` is counted apart from every verdict: a dispatch whose
envelope reported no dollars still shipped or was still abandoned.
------------------------------------------------------------------ #}
Spend by Card
{% if spend %}
last {{ spend.window_days }}d · cached {{ spend.cache_age_seconds }}s ago
(refreshes every {{ spend.cache_ttl_seconds }}s)
{% endif %}
JSON API ↗
{% if not spend %}
Unavailable.
The spend panel did not load. This is not a statement that nothing was spent.
{% else %}
{% if spend.error %}
Served from cache — the last refresh failed: {{ spend.error }}
{% endif %}
{% if spend.state != "measured" %}
Unmeasurable — {{ spend.headline }}.
{% if spend.reason == "no_attributed_rows" %}
No dispatch in the last {{ spend.window_days }} days recorded a cost envelope against a
card. That is an absence of attribution, not an absence of spending — so this
section reads as unmeasurable rather than as a dollar figure of zero.
{% elif spend.reason == "ledger_unreadable" %}
The agent_token_usage ledger could not be read, so nothing here was measured.
{% else %}
The survey could not be produced ({{ spend.reason }}). A panel that could not run is never
a panel that found nothing.
{% endif %}
{% endif %}
{# The KPI row and the outcome table render in BOTH states (xrv-cost-06).
`shape()` already hands over all five outcomes with every figure None
(`_empty_verdicts`) and the API already serves them, so hiding them here
made the JSON and the page disagree about what this panel carries. A
verdict absent from the table is indistinguishable from one that measured
zero -- and that argument applies MOST on the board where nothing was
measured. Nothing can be drawn as money: every None branch below renders
an em-dash, so an unmeasured panel contains no dollar and no percentage.
The CAPTIONS do not carry over, though. `total_cost_usd is none` means
"dispatches ran and none reported a price" when measured and "nothing was
attributed at all" when not, and those are different findings with
different fixes -- the panel counts `unpriced_dispatches` separately for
exactly that reason. #}
{% set unmeasured = spend.state != "measured" %}
{# ONE spelling, pinned to `spend.UNATTRIBUTED_CAPTION` by
tests/cache_savings/test_spend_panel.py so the two cannot drift. #}
{% set unattributed = "nothing in this window was attributed to a card" %}
{% if spend.total_cost_usd is none %}
—
Attributed Spend
{{ unattributed if unmeasured else "no dispatch reported a price" }}
{% else %}
${{ "%.2f"|format(spend.total_cost_usd) }}
Attributed Spend
{{ spend.tasks }} card(s), last {{ spend.window_days }}d
{% endif %}
{# None means nothing was MEASURED — an all-unmeasurable window is not
"0% shipped", and a board with no rows is not "100% shipped". #}
{% if spend.shipped_cost_share_pct is none %}
—
Spend That Shipped
not measured — {{ unattributed if unmeasured else "no judged spend in window" }}
{% else %}
{{ spend.shipped_cost_share_pct }}%
Spend That Shipped
share of the measured spend
{% endif %}
{{ spend.unpriced_dispatches if spend.unpriced_dispatches is not none else "—" }}
Unpriced Dispatches
{{ unattributed if unmeasured else "ran, but reported no dollars" }}
{{ spend.unmeasurable_tasks if spend.unmeasurable_tasks is not none else "—" }}
Unmeasurable Cards
{{ unattributed if unmeasured else "outcome could not be established" }}
| Outcome |
Cards |
Dispatches |
Cost |
Share |
What it means |
{% for row in spend.by_verdict %}
| {{ row.label }} |
{{ row.tasks if row.tasks is not none else "—" }} |
{{ row.dispatches if row.dispatches is not none else "—" }} |
{%- if row.cost_usd is none %}—
{%- else %}${{ "%.4f"|format(row.cost_usd) }}{% endif -%}
|
{{- (row.cost_share_pct ~ "%") if row.cost_share_pct is not none else "—" -}}
|
{{ row.note }} |
{% endfor %}
{% if spend.by_task %}
| Card |
Outcome |
Board status |
Dispatches |
Cost |
Model(s) |
{% for row in spend.by_task %}
| {{ row.task_id }} |
{{ row.label }} |
{{ row.status or "not on board" }} |
{{- row.dispatches -}}
{%- if row.unpriced_dispatches %} ({{ row.unpriced_dispatches }} unpriced){% endif -%}
|
{%- if row.cost_usd is none %}unpriced
{%- else %}${{ "%.4f"|format(row.cost_usd) }}{% endif -%}
|
{{ row.models|join(", ") or "—" }} |
{% endfor %}
{% endif %}
{%- if unmeasured %}
{# "Showing 0 of 0 attributed card(s)" is TRUE for an empty window and
FALSE for a ledger that could not be read -- there, how many cards were
attributed is precisely what is unknown. So the count is withheld and
the table is described instead. #}
No card rows: {{ spend.headline }}. Every figure above is withheld rather than set to zero.
{%- else %}
Showing {{ spend.tasks_shown }} of {{ spend.tasks_total }} attributed card(s), costliest first.
{%- endif %}
Cost is the Claude Code envelope's own total_cost_usd per dispatch, recorded at
reap against agent_token_usage.task_id; the outcome is re-derived from git
(did a commit naming this card land on the default branch) and the board, with the
forge deliberately not consulted so a page render costs no network round-trip.
This spend is not netted against the caching figures above — those
describe calls avoided and input tokens cached, which is a different question and a
different substrate. Re-derive any figure here with
python -m tools.cache_savings.spend --json.
{% endif %}
{% include "includes/iqe_query_widget.html" %}
Pricing model: Anthropic claude-sonnet-4-6 — Input $3/MTok · Output $15/MTok · Cache write $3.75/MTok · Cache read $0.30/MTok.
Context cache savings = read_tokens × ($3.00 − $0.30)/MTok minus write premium.
Response cache savings = avoided_calls × (input_cost + output_cost) per entry.
NIST 800-53: SC-28, AU-12, SA-11.
{% endblock %}