{% extends "base.html" %} {# CUI // SP-CTI — Cache Savings Dashboard (D-CACHE-VIS-1) #} {# NIST 800-53: SC-28 (Protection at Rest), AU-12 (Audit Record Generation), SA-11 (Developer Testing) #} {% block title %}Cache Savings — ICDEV™{% endblock %} {% set iqe_canvas = "cache_savings" %} {% set iqe_api_route = "/api/core/iqe-query" %} {% set iqe_title = "Query Cache Stats" %} {% set iqe_examples = [ {"label": "Hit rate by function", "query": "SELECT function, hit_rate_pct FROM cache.stats ORDER BY hit_rate_pct DESC"}, {"label": "Top functions by hits", "query": "SELECT function, total_hits, avoided_calls FROM cache.stats ORDER BY total_hits DESC LIMIT 10"}, {"label": "Cost savings breakdown", "query": "SELECT function, cost_usd_saved FROM cache.stats ORDER BY cost_usd_saved DESC"}, {"label": "Prefix cache by provider", "query": "SELECT provider, status, cached_share_pct, usd_saved FROM cache.by_provider ORDER BY calls DESC"}, {"label": "Providers with no traffic", "query": "SELECT provider, capability FROM cache.by_provider WHERE status = 'no_data'"}, {"label": "Spend that did not ship", "query": "SELECT task_id, verdict, status, cost_usd FROM cache.spend WHERE verdict != 'shipped'"}, {"label": "Cards with unpriced dispatches", "query": "SELECT task_id, verdict, dispatches, unpriced_dispatches FROM cache.spend WHERE unpriced_dispatches > 0"}, ] %} {% block content %}
Cache Savings {{ "ENABLED" if stats.enabled else "DISABLED" }} backend: {{ stats.backend }} JSON API ↗
{# None means the cache was never ASKED in the window — there is no rate. Rendering 0% there says the cache failed every request it was given, which is the opposite of what happened (cch-obs-03). #} {% if stats.summary.hit_rate_pct is none %}
—
Cache Hit Rate
not measured — no requests in window
{% else %}
{{ stats.summary.hit_rate_pct|round(1) }}%
Cache Hit Rate
{% endif %}
{{ "{:,}".format(stats.summary.total_entries) }}
Cached Entries
{{ "{:,}".format(stats.summary.total_hits) }} total hits
{{ "{:,}".format(stats.summary.cache_read_tokens) }}
Context Cache Reads
{{ "{:,}".format(stats.summary.cache_write_tokens) }} written
${{ "%.4f"|format(stats.summary.total_usd_saved) }}
Estimated Savings (USD)
resp: ${{ "%.4f"|format(stats.summary.resp_cache_usd_saved) }} + ctx: ${{ "%.4f"|format(stats.summary.context_cache_usd_saved) }}
{% if stats.by_function %}
By Function
{% for fn in stats.by_function %} {% endfor %}
Function Entries Hit Rate Avoided Calls Read Tokens Cost Saved
{{ fn.function }} {{ "{:,}".format(fn.total_entries) }} {% if fn.hit_rate_pct is none %} — {% else %} {{ fn.hit_rate_pct }}% {% endif %} {{ "{:,}".format(fn.avoided_calls) }} {{ "{:,}".format(fn.cache_read_tokens) }} ${{ "%.4f"|format(fn.cost_usd_saved) }}
{% else %}
No cached entries yet. Cache activates after the first LLM invocation on an enabled function.
{% endif %} {# ------------------------------------------------------------------ Per-provider PREFIX cache effectiveness (cch-obs-01). Everything above this point describes the RESPONSE cache — calls avoided outright — which is legitimately one number. This section describes PREFIX caching, which is not: the providers do different things, so one rate over the mix is a blur rather than a summary. The colour rules are the whole point. Only `caching` and `no_cache_hits` are on a red/amber/green scale, because only they measured something. `no data` and `not reported` are neutral and say why — a provider with no traffic and a provider that reports no counters are not failing caches. ------------------------------------------------------------------ #}
Prefix Cache by Provider {% if by_provider and by_provider.measurable %} last {{ by_provider.window_days }}d · {{ by_provider.window_start }} to {{ by_provider.window_end }} {% endif %} JSON API ↗
{% if not by_provider or not by_provider.measurable %}
Unmeasurable. {{ by_provider.unmeasurable_reason if by_provider else "the per-provider view did not load." }}
This is deliberately not shown as 0% on every provider — a database with no operating history and a platform where caching is broken would look identical.
{% else %}
{% for p in by_provider.providers %} {% set measured = p.status in ['caching', 'no_cache_hits'] %} {% endfor %}
Provider Caching Status Calls Cached In Uncached In Cached Share Trend Saved
{{ p.provider }} {{ p.capability|replace('_', ' ') }} {{ p.status_label }} {{ "{:,}".format(p.calls) }} {{ "{:,}".format(p.cached_input_tokens) }} {{ "{:,}".format(p.uncached_input_tokens) }} {%- if p.cached_share_pct is none %} {# NOT 0%. There is no rate to report here, and printing one would be the exact conflation this card was opened to remove. #} — {%- else %} {{ p.cached_share_pct }}% {%- endif %} {%- if p.trend.direction == 'improved' %} ▲ {{ p.trend.delta_pct_points }} pp {%- elif p.trend.direction == 'worsened' %} ▼ {{ p.trend.delta_pct_points }} pp {%- elif p.trend.direction == 'flat' %} flat {%- else %} no baseline {%- endif %} {%- if p.usd_saved is none %} {# A local provider has no bill, so it has no dollars to save. $0.00 here would read as "caching failed" for a cache that is working fine and simply is not billed. Latency instead. #} {# A local provider's latency is its only observable effect — but only when it was actually called. "0 ms avg" for a provider with no traffic is the same fabricated zero this table exists to remove, one column over. #} {%- if p.usd_basis == 'local' and p.calls > 0 %}{{ p.avg_latency_ms|round(0)|int }} ms avg {%- elif p.usd_basis == 'local' %}— {%- else %}not priced{% endif -%} {%- else %} ${{ "%.4f"|format(p.usd_saved) }} {%- endif %}
{{ by_provider.totals.providers_caching }} caching · {{ by_provider.totals.providers_no_cache_hits }} measured zero · {{ by_provider.totals.providers_unreported }} not reported · {{ by_provider.totals.providers_no_data }} no data. Total saved ${{ "%.4f"|format(by_provider.totals.usd_saved_total or 0) }} — {{ by_provider.totals.usd_saved_basis }}.
There is deliberately no combined hit rate on this table. Providers disagree on whether reported input tokens already include the cached ones (Anthropic and Bedrock report them separately; OpenAI and Azure report cached tokens as a subset), so one average over the mix would double-count every cached token from half of them. Claims live in args/cache_effectiveness.yaml.
{% endif %}
{# ------------------------------------------------------------------ SPEND BY CARD (xrv-cost-04) A different question from everything above it. The tables above ask what CACHING saved; this asks what the board SPENT and whether that spend shipped. They are never added together and never share a rate. Every empty state is rendered IN WORDS. $0.00 on a cost panel reads as "this was free", which is the one thing an unattributed board must never be mistaken for — the cards may well have cost money, nothing attributed it. `unpriced` is counted apart from every verdict: a dispatch whose envelope reported no dollars still shipped or was still abandoned. ------------------------------------------------------------------ #}
Spend by Card {% if spend %} last {{ spend.window_days }}d · cached {{ spend.cache_age_seconds }}s ago (refreshes every {{ spend.cache_ttl_seconds }}s) {% endif %} JSON API ↗
{% if not spend %}
Unavailable. The spend panel did not load. This is not a statement that nothing was spent.
{% else %} {% if spend.error %}
Served from cache — the last refresh failed: {{ spend.error }}
{% endif %} {% if spend.state != "measured" %}
Unmeasurable — {{ spend.headline }}.
{% if spend.reason == "no_attributed_rows" %} No dispatch in the last {{ spend.window_days }} days recorded a cost envelope against a card. That is an absence of attribution, not an absence of spending — so this section reads as unmeasurable rather than as a dollar figure of zero. {% elif spend.reason == "ledger_unreadable" %} The agent_token_usage ledger could not be read, so nothing here was measured. {% else %} The survey could not be produced ({{ spend.reason }}). A panel that could not run is never a panel that found nothing. {% endif %}
{% endif %} {# The KPI row and the outcome table render in BOTH states (xrv-cost-06). `shape()` already hands over all five outcomes with every figure None (`_empty_verdicts`) and the API already serves them, so hiding them here made the JSON and the page disagree about what this panel carries. A verdict absent from the table is indistinguishable from one that measured zero -- and that argument applies MOST on the board where nothing was measured. Nothing can be drawn as money: every None branch below renders an em-dash, so an unmeasured panel contains no dollar and no percentage. The CAPTIONS do not carry over, though. `total_cost_usd is none` means "dispatches ran and none reported a price" when measured and "nothing was attributed at all" when not, and those are different findings with different fixes -- the panel counts `unpriced_dispatches` separately for exactly that reason. #} {% set unmeasured = spend.state != "measured" %} {# ONE spelling, pinned to `spend.UNATTRIBUTED_CAPTION` by tests/cache_savings/test_spend_panel.py so the two cannot drift. #} {% set unattributed = "nothing in this window was attributed to a card" %}
{% if spend.total_cost_usd is none %}
—
Attributed Spend
{{ unattributed if unmeasured else "no dispatch reported a price" }}
{% else %}
${{ "%.2f"|format(spend.total_cost_usd) }}
Attributed Spend
{{ spend.tasks }} card(s), last {{ spend.window_days }}d
{% endif %}
{# None means nothing was MEASURED — an all-unmeasurable window is not "0% shipped", and a board with no rows is not "100% shipped". #} {% if spend.shipped_cost_share_pct is none %}
—
Spend That Shipped
not measured — {{ unattributed if unmeasured else "no judged spend in window" }}
{% else %}
{{ spend.shipped_cost_share_pct }}%
Spend That Shipped
share of the measured spend
{% endif %}
{{ spend.unpriced_dispatches if spend.unpriced_dispatches is not none else "—" }}
Unpriced Dispatches
{{ unattributed if unmeasured else "ran, but reported no dollars" }}
{{ spend.unmeasurable_tasks if spend.unmeasurable_tasks is not none else "—" }}
Unmeasurable Cards
{{ unattributed if unmeasured else "outcome could not be established" }}
{% for row in spend.by_verdict %} {% endfor %}
Outcome Cards Dispatches Cost Share What it means
{{ row.label }} {{ row.tasks if row.tasks is not none else "—" }} {{ row.dispatches if row.dispatches is not none else "—" }} {%- if row.cost_usd is none %}— {%- else %}${{ "%.4f"|format(row.cost_usd) }}{% endif -%} {{- (row.cost_share_pct ~ "%") if row.cost_share_pct is not none else "—" -}} {{ row.note }}
{% if spend.by_task %}
{% for row in spend.by_task %} {% endfor %}
Card Outcome Board status Dispatches Cost Model(s)
{{ row.task_id }} {{ row.label }} {{ row.status or "not on board" }} {{- row.dispatches -}} {%- if row.unpriced_dispatches %} ({{ row.unpriced_dispatches }} unpriced){% endif -%} {%- if row.cost_usd is none %}unpriced {%- else %}${{ "%.4f"|format(row.cost_usd) }}{% endif -%} {{ row.models|join(", ") or "—" }}
{% endif %}
{%- if unmeasured %} {# "Showing 0 of 0 attributed card(s)" is TRUE for an empty window and FALSE for a ledger that could not be read -- there, how many cards were attributed is precisely what is unknown. So the count is withheld and the table is described instead. #} No card rows: {{ spend.headline }}. Every figure above is withheld rather than set to zero. {%- else %} Showing {{ spend.tasks_shown }} of {{ spend.tasks_total }} attributed card(s), costliest first. {%- endif %} Cost is the Claude Code envelope's own total_cost_usd per dispatch, recorded at reap against agent_token_usage.task_id; the outcome is re-derived from git (did a commit naming this card land on the default branch) and the board, with the forge deliberately not consulted so a page render costs no network round-trip.
This spend is not netted against the caching figures above — those describe calls avoided and input tokens cached, which is a different question and a different substrate. Re-derive any figure here with python -m tools.cache_savings.spend --json.
{% endif %}
{% include "includes/iqe_query_widget.html" %}
Pricing model: Anthropic claude-sonnet-4-6 — Input $3/MTok · Output $15/MTok · Cache write $3.75/MTok · Cache read $0.30/MTok. Context cache savings = read_tokens × ($3.00 − $0.30)/MTok minus write premium. Response cache savings = avoided_calls × (input_cost + output_cost) per entry. NIST 800-53: SC-28, AU-12, SA-11.
{% endblock %}