compression with a quality contract

Changelog

Every notable change to Distil, newest first — generated from CHANGELOG.md, so this page can never drift from the source of truth.

Unreleased

1.54.0 — 2026-09-25 — measured, enforced, verifiable

In short

  • Added — MCP tool search stays on. distil wrap / --always-on set ENABLE_TOOL_SEARCH, so Claude Code defers MCP tool definitions behind the proxy (verified live). distil discover names connectors you pay for and never call.
  • Added — cold-point recompression. Older tool output is re-stubbed only when the prompt cache has certainly expired (--no-cold-point to opt out).
  • Added — one risk budget and a drift alarm that acts. A proven decision-drift breach holds the proxy at lossless-only; release with distil reset --drift-guard.
  • Added — sealed receipt segments with Merkle checkpoints. One receipt can be proven to an auditor without the whole log (distil receipts --prove/--check-proof).
  • Added — agent presets fixed and contract-tested (Kilo, aider, opencode, qwen, openhands, grok, kimi; Cline CLI).
  • Fixed — receipts no longer re-read the whole chain per request; an in-window prefix-cache break; Anthropic compaction / OpenAI encrypted items are now counted; discover --since keeps recently-failed sessions; auto output-shaping can no longer keep itself on; Windows fixes.
  • Research — where the live bill goes (#189); measured decisions not to compress tool schemas (ADR 0012) or add per-command shell profiles (ADR 0015).

The full story, change by change:

Three threads, and the same shape keeps recurring below. The first is the re-read delta measured against its own written contract: rules stated in an ADR and not implemented in the path that runs them. The second is the exposed surfaces measured against the guards distil already applies elsewhere: a body the proxy refuses and the gateway read as empty, a tenant label the client-supplied header validates and the identity claim did not, a socket timeout the proxy sets and the component you actually bind to a network did not. The third is distil wrap measured against the agents it claims to reach: a preset verified from someone else's documentation rather than guessed at, and a config patched where that agent will actually look for it rather than where distil assumed. None of them is a new capability. All are the distance between what the documentation promises and what the code does, which is the one kind of defect a soak cannot be relied on to surface.

Windows: two quarantines in one clock tick, and a clock constant that is not there

  • A second unreadable drift.json quarantined within one clock tick of the first got the same drift.json.corrupt-* name. On Windows, where the clock ticks every ~15.6 ms, the copy then failed outright and the second corruption was never set aside; elsewhere the copy fallback overwrote the first archive. A clashing name now takes a -N suffix, and the no-hard-link fallback creates its copy exclusively, so no archive is ever overwritten. The cold-point clock picker looks up CLOCK_MONOTONIC with getattr, like CLOCK_BOOTTIME already was, so an interpreter without it (every Windows Python) falls back to time.monotonic instead of raising. The sh escape hatch's owner-copy step tolerates a Python with no os.chown. Tests that redirected HOME now redirect USERPROFILE too. Without it, the Windows CI runner's own ~/.claude/settings.json was being wired. The test suite now gives every test its own HOME and USERPROFILE (and unsets CLAUDE_CONFIG_DIR, XDG_CONFIG_HOME, CLINE_DATA_DIR), so no test can write a real agent config on any OS even if it forgets to redirect the home. A real_home marker opts a read-only resolution test out, and a guard test fails if any config resolver points outside the sandbox.

Tool definitions: measured, deliberately left alone

  • benchmarks/tool_schema_share.py measures what request tools definitions cost on real traffic, using the per-request records distil wrap already writes. It spreads each request's tool tokens over the provider's cache-read, cache-write and uncached counts in prefix order and prices them. With --tools it also measures how much of a real tools array a lossless compactor could remove. The first result is committed at benchmarks/results/2026-09-24/tool_schema_share.json. Tools are a large share of the bill, but almost all of it is cache reads, and the lossless headroom is small. A perfect compactor would save about 1.1% of billed cost, below the bar for a default-on rewrite of the cache prefix. ADR 0012 records why distil does not do this and what would change the decision. No runtime behaviour changes.

The same shape keeps recurring below. The first half is the re-read delta measured against its own written contract: rules stated in an ADR and not implemented in the path that runs them. The second half is the exposed surfaces measured against the guards distil already applies elsewhere: a body the proxy refuses and the gateway read as empty, a tenant label the client-supplied header validates and the identity claim did not, a socket timeout the proxy sets and the component you actually bind to a network did not. Neither half is a new capability. Both are the distance between what the documentation promises and what the code does, which is the one kind of defect a soak cannot be relied on to surface.

1.53.0rc1 shipped the re-read delta with its contract written out in six rules, an ADR and a changelog entry. A read of the request path against that contract found two of the rules stated and not implemented, and six older defects on the paths the new transform now runs more of. A cross-audit of the result then found one more, introduced by this branch: the fix that moved the session lock's sidecar out of the user's config directory made the lock depend on DISTIL_HOME, which is to say it stopped being a lock. That is reverted below, litter and all, with the reasoning written down this time. Nothing here is a new capability. Each one is the difference between what the documentation promises and what the code does, which is the only kind of bug a soak cannot be relied on to surface — the shapes below are invisible under distil wrap on Claude Code, and that is the only traffic the soak has.

Alongside them runs the same measurement turned outward. Every piece of statistical machinery in this repo already worked; none of it was ever shown to the person whose traffic it was measuring. That is not a gap in rigor, it is rigor that stayed in the library while the user got a savings number.

Anthropic server-side compaction: census gap fixed, passthrough pinned by contract tests

Anthropic's Messages API now compacts context server-side (beta compact-2026-01-12 / compact-2026-09-04) and returns a compaction content block whose signature the provider re-validates on replay: alter it, move it, or even re-encode it losslessly, and the next request 400s with compaction_signature_invalid. A new contract test suite (tests/test_compaction_passthrough.py, 22 cases) pins that distil's whole Anthropic path — compress_messages (digest/recency/provenance/rereaddelta), the SDK wrap() adapter, the proxy (context_management field + anthropic-beta header), and the streaming splice — never touches that block, both non-streaming and streaming, including when the tool_results around it ARE digested.

Passthrough was already safe: the dispatch that makes it true predates this work (an unknown block type falls through untouched, the same guard that already protected thinking/redacted_thinking), and 18 of the 22 cases pass with no code change at all — they exist now as a contract so it stays true. The actual bug the suite found was a census gap: a compaction block's billed tokens were invisible to the eligibility census, the exact blind spot thinking_billed was added to close for extended thinking. Fixed by generalising that guard from an allowlist of two type strings to "provider-signed and opaque" (compaction, or any future block carrying a signature) so a cost distil cannot reduce is not also one it hides from the savings percentage — narrowed to exclude blocks with their own dedicated handling (tool_result, text, tool_use, image) so a stray signature key on one of those can't shadow its compression or its census. dissect.py's protected/missed-opportunity split now recognises compaction_billed and signed_block_billed alongside thinking_billed, so a compaction- or thinking-heavy session reads as policy holding rather than a broken gate.

The OpenAI adapter had no guarantee the Anthropic path had just proven

A sibling audit proved Anthropic's server-side compaction block — the one the provider re-validates by signature on the next turn — never gets touched, moved, or re-encoded by compress_messages. The OpenAI Responses API has the exact same class of surface and no equivalent proof had ever been written: a reasoning item's encrypted_content (stateless mode / Zero Data Retention) and a compaction item — precisely what POST/v1/responses/compact returns, and what OpenAI's own docs say not to prune — must reach the next request byte-identical or the provider cannot re-derive state it never got back. Twenty contract tests (tests/test_openai_opaque_passthrough.py) found the passthrough itself was already correct: _compress_response_item dispatches by item["type"], and neither item type is on its digest path, so both survive every mode — verbatim, digest, recency, the distil_expand re-query, and the buffered-stream re-emit — as the same object, never a copy. What was missing was the accounting: three tests failed before any fix, because these items' tokens landed in neither count_responses_tokens (by design — they join assistant_text/function_call outside the compressible-zone baseline, not a bug) nor the eligibility census, so a request could show near-zero savings with the real explanation — thousands of billed, uncompressible reasoning/compaction tokens — invisible. Same shape as the Anthropic gap, same fix: _census_opaque_response_item attributes encrypted_content (and any summary/text/content a future variant carries) to reasoning_billed / compaction_billed, generalised on the presence of encrypted_content rather than an allowlist of type strings — a future opaque item type is safe by construction. /v1/responses/compact itself was already never touched: it fails the is_responses_path regex (a distinct endpoint, not a query-string variant) and falls through to the byte-for-byte _passthrough relay, pinned here by test rather than by reading the regex and hoping. context_management and previous_response_id were already forwarded unchanged (a request body spread that never drops an unrecognised key), also now pinned. A review pass on that fix found the opaque-item guard was broader than it needed to be: "encrypted_content" in item alone, without also excluding the known compressible types, could in principle let a stray or future encrypted_content key on a message/ function_call_output/function_call item shadow its own real handling. Narrowed to exclude those three, with a test pinning a message carrying a decoy encrypted_content key still compresses and censuses as user_text, not as an opaque passthrough. The census gap's mirror in dissect's eligibility report is also closed: reasoning_billed, compaction_billed, and signed_item_billed are now in _ELIGIBILITY_LABEL and _PROTECTED_REASONS, so a reasoning-heavy session reads as "the design holding," the same verdict Anthropic's thinking/compaction census buckets already earn — not as a missed compression opportunity. And because count_responses_tokens deliberately does not count these items (documented at the function and in cache-contract.html clause (g)), the eligibility census total can legitimately exceed x-distil-compressible-tokens on a Responses session; both the docstring and the doc now say so, and both the census tokens and the report's opaque-bucket labels are marked approximate — a heuristic count of base64 ciphertext, not the provider's real billed reasoning-token count.

The drift alarm stops the compression it catches, against the one budget everything else reads

Two gaps, one shape. Before this, the anytime-valid drift e-process could print BREACHED at wrap exit while the proxy carried on compressing, because nothing on the request path imported distil/drift.py. The budget it bet against was also one of three thresholds nobody coordinated: the drift alarm's α, the conformal certificate's α, and certify's TOST margin, each a literal in its own file. On the maintainer's own shadow ledger (read-only, 2026-09-24) that produced a real contradiction: the proof ledger printed certified decision-change budget: intact directly above a conformal bound that was over the budget.

  • One risk budget. distil.conformal now owns BUDGET_ALPHA (≤5% decision change), BUDGET_DELTA (at 95% confidence) and CERT_MARGIN (the 2 pp TOST margin). The conformal certificate, certify-trajectories, calibrate, the drift e-process, the proof ledger's budget and risk lines, the proxy guard, and every CLI and library default all read them at call time. No values changed. The TOST margin and the budget measure the same estimand, and the margin is deliberately stricter, so a point that just certified does not trip the live alarm on noise. A test now pins CERT_MARGIN ≤ BUDGET_ALPHA, and another changes BUDGET_ALPHA and checks that the certificate, drift, risk and guard verdicts all move with it.
  • intact is earned. When the e-process has not tripped, that means no breach has been proven. It does not mean "within budget". The budget line now prints intact only when the bound next to it is inside the budget. Otherwise it prints unproven. The risk line now says within or ABOVE the budget.
  • One e-process, written only by the proxies that feed it. drift.json is the e-process. Each proxy folds the paired shadow verdict it just produced into that file, under a file lock, from the shadow thread (drift.fold). Loading inside the lock means a restart, a hot-swap worker and a second wrap all continue the same capital, so no row is bet twice and no restart takes a fresh look from stale capital. The new tests show the restart path produces the exact same capital as one long-lived process, and that the null false-alarm rate across simulated restarts stays at δ. Wrap exit, distil stats and the status line only read the file; a reporting command never writes it. Neither does distil doctor: its proxy self-test runs a read-only guard, with no migration and no watcher thread. An existing old-format drift.json is migrated once, on the first proxy start after this upgrade, by rebuilding it from shadow.jsonl in file order. That is the same fold the exit summary used to do. A missing drift.json is folded from shadow.jsonl once, but only on a first-ever start: a fresh install, or an upgrade from a version that never wrote the file. Starting those at zero would throw away harm evidence the machine had already measured. If a release archive (drift.json.reset-*) sits beside the missing file, the user has released before. That case starts fresh and never re-folds, or a release whose fresh state was deleted or never written would re-trip on the very rows it released. A quarantined .corrupt-* file does not count as a release. If both the state file and every release archive are deleted, the machine looks like a first install and re-bootstraps; that is accepted.
  • A state file nobody can read is held, not reset. A zero-length, garbage or wrongly-typed drift.json used to load as a fresh monitor, and the next fold then overwrote it, so a recorded breach could vanish. Now:
    • It counts as held. The first writer copies it to drift.json.corrupt-<time>-<ns> (a hard link, or a copy when linking fails), then atomically writes a held state over the original that records the copy's name. The file is copied rather than moved so there is no moment when drift.json is missing; a missing file would read as released. If the held write fails, the corrupt file stays in place and is still held. If the copy fails, the original is never overwritten. A second corruption gets its own name.
    • Every surface says HELD — the drift state file was unreadable …, followed by the release command.
    • Writes fsync before the atomic rename.
    • Known limits, documented rather than built:
      • When the advisory lock can't be taken, locking fails open. Two processes crossing the threshold together can then each write a trip receipt; the hold itself is unaffected.
      • .reset-* and .corrupt-* archives are never pruned, because they are evidence. They accumulate one per release or corruption.
  • The alarm acts. On a breach, the next request is served lossless-only: Tier-0, no digest, no output shaping. The response carries x-distil-mode: lossless-only and x-distil-drift-guard: held. distil_expand stays injected, so stubs already in the history stay recoverable and the cached tools prefix keeps its shape. The hot-path cost is one attribute read, with no per-request file I/O. Held requests are receipted as byte-reversible (lossless-only is Tier-0). A request that issued digest handles never is, whatever its mode. This also corrects --lossless-only, which the receipt chain used to mislabel as not reversible.
  • It persists globally, and every process sees it. A trip is part of drift.json and adds one mode: drift-trip event receipt (zero counts, reversible: false) to the chain. It is written inside the same lock, so two processes crossing together produce one receipt. The hold is global rather than per-session: the e-process and budget already are, and a per-session hold would let the next distil wrap resume lossy compression right after a certified breach. A long-running proxy, such as the launch agent, stats the file every 30s from a daemon thread. It picks up another process's trip, and a release, without a restart. The async proxy is already Tier-0, so a trip there turns off output shaping. The multi-tenant gateway runs no shadow and is deliberately exempt. One tenant's evidence must not hold every tenant; see ADR 0016.
  • Visible where you look, releasable without collateral. While a hold is on, the status line shows ⚠ drift hold · distil reset --drift-guard. The wrap exit summary and distil stats say BREACHED … compression held at lossless-only, followed by the same command. distil reset --drift-guard archives only the drift state and leaves a fresh one, so the next start does not re-fold the same rows into the same breach. Savings and shadow stats are untouched, and running proxies resume within 30s. If the release cannot archive the old state or write the fresh one, it says so on stderr and exits 1; it never claims a release it did not make. distil reset--shadow still releases the hold as well.
  • Fail-open, on by default. A guard that raises, or cannot load its state, serves the request exactly as configured. DISTIL_NO_DRIFT_GUARD=1 opts out of the hold. The alarm still trips, and the proof line then says compression was not held.
  • Upgrading with an e-process that has already tripped. If an earlier version left a ~/.distil/drift.json that says BREACHED, the proxy starts held at lossless-only. That is the alarm doing its job on evidence you already had. An old-format file that has not tripped is rebuilt from shadow.jsonl, and it holds if that evidence crosses the budget. With no drift.json at all (a version that never wrote one), the existing shadow evidence is folded once on the first start, and it holds if it crosses the budget. After distil calibrate, release a hold with distil reset --drift-guard.
  • Do not run an older distil side by side. An older build still installed next to this one, such as a second venv or a pinned launch agent, rewrites drift.json in the old format at its own wrap exit. The next proxy start from this build then sees an old-format file and rebuilds the e-process from shadow.jsonl again, from K_0 = 1. Each alternation between the two versions is another look at the same rows, so the one-e-process guarantee only holds when a single version writes the file.

Old tool output shrinks on the turn the cache has already expired

A tool result is billed on every later turn at the cache-read rate. Shrinking it after first sight used to bust the cache, which is what the August 2x incident did. The one exception is a turn that arrives after the provider's cache entry has expired: the whole prefix is re-written then anyway. On such a turn distil now replaces older tool output with a stub that distil_expand can recover (<<distil evicted older tool output (N lines);distil_expand handle=H recovers it>>). It then forwards those same stubs on every later turn, so the smaller prefix is what gets cached for the rest of the session (ADR 0014, distil/coldpoint.py).

  • Expiry has to be certain from distil's own observation. That means a lineage known to this process, nothing of it in flight (streams, expand re-queries and shadow replays count), and a monotonic gap since distil last finished forwarding it that exceeds the longest TTL the lineage asked for (5 min, or 1 h for ttl: "1h") plus 60 s. An unparseable TTL never evicts. First-seen, restarted and hot-swapped lineages do nothing.
  • Two conversations under one lineage key never evict. Every request has to extend the previous one. Parallel subagents, forks and rewinds fail that test, and the lineage turns ambiguous for good.
  • Byte-stable after the cold point. Stubs are a pure function of content, the evicted tool_use_id set only grows, and prefix replay holds the result. A proxy test drives first-seen → warm → 20-minute gap → four warm turns and asserts that the cold turn's messages are forwarded byte-identical on every later turn.
  • Never evicted: the freshest tool turns, exact-quote results (Edit targets, shell reads, distil_expand results; if an Edit comes to depend on an evicted block later, the exemption wins), text an Edit already quotes, learned-keep content, expanded handles, and blocks under 128 tokens.
  • Where it runs: Anthropic Messages on distil proxy / distil wrap, only where the recoverable digest already runs. Lossless-only, verbatim and --session-delta are untouched. Opt out with --no-cold-point or DISTIL_COLD_POINT=0. The gateway and OpenAI are documented TODOs in the ADR.
  • Accounting uses the existing path. Evicted tokens are part of tokens_saved. The census gets a new tool_result_evicted bucket, which dissect labels. Handles are in the receipt (mode digest). x-distil-cold / x-distil-cold-evicted go on the response, and cold / cold_evicted go in the session ledger.
  • The set survives a hot-swap or restart within one wrap session. Evicted ids are persisted content-free to $DISTIL_HOME/coldpoint.json (0600, atomic, at most 512 lineages × 2048 ids, written only when a set grows, parsed once per change rather than per request), so a hot-swap re-applies the same stubs instead of un-evicting a warm prefix. A fresh distil wrap or a claude --resume in a new terminal is a new session id and a new lineage, so it still pays one rewrite. The lineage is keyed on the static API key, not the OAuth bearer, so a token refresh does not fork it. The idle clock counts system sleep (CLOCK_MONOTONIC on macOS, CLOCK_BOOTTIME on Linux).
  • Not yet measured live. The gate is an rc soak plus a live A/B on distil cache read/write totals, and neither has been run. The ADR names two soak gates: the distribution of cold reasons, and zero cold turns that still read history from cache.

An Edit no read could have carried re-wrote the cached prefix

The quote guard checks that every Edit's old_string still occurs in the payload about to be forwarded, and on a miss re-compresses with the exact-quote exemption widened. It then forwarded the widened pass whether or not the pass found the quote. On Claude Code it almost never can. A multi-line quote does not occur in line-numbered Read output, and text the agent Write-ed itself was never in a tool result at all. So the first such Edit in a session turned every re-read stub and superseded read already in the provider's cached prefix back into verbatim text. That was a whole-prefix re-write at the write rate, inside the cache window, for nothing. And because the history only grows, it kept the class off for the rest of the session. This is the "provenance re-expand" the live-savings research counted among in-window breaks, and it is a different bug from #184's expand re-digest. The widened pass is now forwarded only when it rescues at least one quote and loses none the narrow pass kept, on both the Messages and the Responses paths (_widen_rescued). The lost quotes are compared as sets, not counts, so the choice does not assume the widened pass is monotone. The miss is still booked in quotes. Replay now calls a walk "held" only when it reached the end of the previous turn on both the client's list and the forwarded list, so a stop caused by distil's own forwarded list can't be labelled "held". Replayed offline over 12 local Claude Code transcripts (4,658 append-only requests), the widened pass ran on 2,625 requests and rescued a quote on none. The same replay found one in-window break before the change (176,944 bytes of prefix) and none after it. Prefix replay now also records why it stopped: replay_stop in the request ledger, x-distil-replay-stop on the response. The values are held, client, distil, marker, cold and untracked. Before this, an in-window cache write could not be blamed on either side. On the maintainer's ledger, $25.26 of the $35.83 in-window break cost (1.28% of billed) sits on turns where replay diverged with no record of whose rewrite it was. benchmarks/in_window_prefix_breaks.py produces both halves as content-free JSON under benchmarks/results/2026-09-24/. docs/cache-contract.html gains clause (g).

The freshest read the agent asked for came back as a pointer

ADR 0010 rule 0 says the newest tool output is never elided, for the reason the recency carve-out exists at all: a re-read the agent has just issued is precisely the output it reasons over to choose its next action. The planner honoured it — which blocks may serve as bases never depended on the sliding window — and the application did not. So an agent that re-read a file to find out whether its edit had landed was handed «distil-reread handle=…» lines 1-80 of this result are byte-identical to… and a round trip through distil_expand to get the answer back.

The gate is now where the rule always said it was: the plan is still computed from the prefix, and only its application consults recency. That is one line of behaviour and a much larger change to what the tests assert — tests/test_reread_delta.py had no test for this rule at all, and every fixture in it ended on the re-read, which is exactly the position the rule protects. The distil validate re-read cases had the same shape and would have gone quietly green over a transform that no longer ran; they now act on what they read before the payload is measured, as a real session does.

Why a week of soak could not find it. The recency window is anchored to the client's last cache_control breakpoint, so when the client pins its newest block — Claude Code's shape — the window is empty and this gate is a no-op. The defect only reached clients that send no marker: plain SDK callers, third-party agents on /v1/messages, and the opening turns of every session before a marker lands. On the codebench corpus the unmarked row moves from 10.3% to 7.1% tokens, which is the number the 1.53.0 entry below already quotes; the marked row that bills is unchanged at 41.7% and 49.1%.

A CRLF read and an LF read of one file compared equal

Rule 5 says lines carry their terminators into the comparison, and gives the reason: stripping them calls a CRLF read and an LF read of one file identical, so eliding the only LF copy while the base holds CRLF bytes makes the stub's claim of byte-identity false. The planner stripped them. str.splitlines() also folds a form feed, a next line and a line separator, so a form-feed-separated read matched a newline-separated one too.

The application side had always sliced with keepends=True, so the two halves of the transform agreed about which lines by coincidence and disagreed about what a line is. They now use the same split, and the stub's claim is true of bytes rather than of characters. Matches get strictly rarer — a run reaching end-of-file in the target but not in the base gives up its last line — which is the safe direction. The exact-quote guard caught the consequence, but only from the turn an Edit with a literal old_string appeared; before that the model composed its quote from the wrong-terminator base.

A recovered block was folded back into the stub it had just escaped

distil_expand is the recoverable half of the product: a digested block keeps a handle, the agent asks for the handle, and the original comes back. Over the MCP server that recovery arrives on the next request as an ordinary tool_result — the client ran a tool and reported what it returned — and the exact-quote exemption had no case for it. So the adapter digested it like any other tool output, and because a digest handle is sha256[:8] of the content, identical bytes produced the identical handle. The agent expanded 4b7055a5, received << +116 lines, handle=4b7055a5 >>, and had no move left that could reach the detail. Observed live.

The exemption now runs before the name test, on every adapter at once, because all three read the same classifier: a tool_result whose paired tool_use is the expand tool is kept byte-exact whatever its age, under its own census bucket (tool_result_expand_recovered) so distil dissect prices it apart from the other two. Both namings are matched — the proxy injects its own tool as distil_expand, and Claude Code namespaces the MCP server's copy as mcp__<serverkey>__distil_expand with a key the user chooses, so the suffix form counts too. The proxy keeps injecting its tool when the MCP one is present; two routes to the same content is not a problem, and this exemption is what makes the second one safe. Matching on the name is the narrow version deliberately: the general fix compares the paired call's input.handle against the handle about to be emitted and needs no name at all, and the comment at the call site says so for the first client that renames the tool instead of namespacing it.

The streaming path resolved expansions and reported none

x-distil-expanded and the receipt's expanded_handles were wired into the buffered expand loop only. The streaming splice resolves handles too — that is the whole point of streamexpand, which intercepts the call mid-stream so recoverable digest costs no time-to-first-token — but it reported nothing, so quality.expand_resolved_requests in distil dissect read 0 on every streaming session no matter how much the agent pulled back, and expansion_regret had no handles to attribute. The feature whose entire claim is that it can prove recovery happened was, on the default path, unable to.

Both branches now share one collection and one reporting step. The wire header cannot follow on the streaming path, and the reason is worth stating rather than working around: response headers are flushed with the first upstream frame, which is before any expand call exists. The receipt is written after the stream closes, and the receipt is what dissect reads.

The same fix deletes an anomaly. "Every request took the streaming pass-through — distil_expand calls could never be intercepted" was written when the streaming path genuinely could not intercept; since the splice landed, its premise is false, and all it detected was a session where the agent never asked to expand — the ordinary healthy case, on which it fired every time. Narrowing it to what it means to warn about, an escaped tool call, is not possible: nothing in a receipt observes one. An unfalsifiable check is worse than no check, because it teaches the reader to skip the anomaly list. Restore it if a receipt ever records whether the expand tool was injected.

The turn a stub first appeared paid to re-bill the whole prefix

_serialize_if_changed exists because the provider's prompt cache matches on exact bytes and json.dumps is not a byte-faithful round trip of what arrived. The streaming expand-intercept path built its upstream body with a bare json.dumps at two places, so the turn a digest handle first entered the conversation — flipping the request onto that path — re-spelled every message with ", " separators and \uXXXX escapes and bought a full cache write for a rewrite the model cannot see. It was also the encoding prefixreplay._wire models, so replay_restored was being measured against bytes that path never sent. Both now go through the same serializer as every other request.

An 8-hex handle could resolve to another block's bytes after a restart

RestoreStore._record refuses a handle that already maps to different text and declines the stub, because expanding it would return the wrong content. The on-disk store made the same check and then returned nothing: it kept the first writer, correctly, but the caller never learned, so the stub went out anyway. In the running process it expanded correctly from memory. After a restart, or from the MCP server in another process, it resolved to the other block. mcp_server.record_restore now reports the collision and _record declines the stub on it, which is what it already did for the in-memory half. A write that merely fails still returns true: persistence is best-effort and the session's own store still answers.

The check itself is an exclusive create rather than p.exists() then write. The old form was check-then-act across processes, which is the only situation this guard is for: two proxies folding the same block could both see "absent", and on a real collision the second would clobber the first — the exact outcome, in the exact concurrent case, that the guard was added to prevent. O_CREAT|O_EXCL lets the filesystem pick the winner, and the loser finds the file already there and compares bytes instead of overwriting. It also means the blob is born 0600 rather than chmod-ed to it afterwards.

The restore store was re-stat'd on every handle it recorded

Recording one handle sorted the whole store by mtime twice — once for the count cap, once for the age sweep — so at the 5,000-file cap a single handle cost up to 10,000 stat() calls. That is the ms/turn ≈1 → ≈21 ADR 0010 attributed to the restore store, and the re-read delta records several handles per turn. The sweep now runs on one listing and at most once per 64 records. Against a 2,000-file store it falls from 16.9 ms to 0.3 ms per recorded handle. The trade is named where it lives: the store may sit up to 63 files above its cap between sweeps.

The TTL is not part of that trade, and the first version of this change made it one. DISTIL_RESTORE_TTL_DAYS is a retention boundary for originals that can hold secrets or PII, not a housekeeping preference, and with the sweep amortized a store that never receives a 64th handle never reaches the trigger — so a quiet machine would keep every expired blob on disk, and keep serving them, for as long as it stayed quiet. Retention enforced only by a schedule that traffic can starve is not retention.

So it is enforced on the read, where nothing can skip it: load_restore compares the blob's mtime against the TTL and treats an expired one as absent, unlinking it on sight and failing open if the unlink loses a race. The sweep keeps its counter and gains a second trigger — once per TTL/24 of elapsed time — so bulk expiry still happens on a low-traffic store without waiting for records that may never come. An expired blob also stops reading as a collision, so it cannot refuse a new stub for the same handle forever.

os.replace onto a contended path is not a permission error on Windows

Every atomic write in this codebase ends the same way: write a temp file beside the target, fsync, swap it in with one os.replace. POSIX rename is atomic against a concurrent rename and against readers, so that is the whole story there. Windows' MoveFileEx has to open the destination, so two writers racing on one target fail each other with ERROR_ACCESS_DENIED or ERROR_SHARING_VIOLATION. The Windows CI gate failed on exactly that, in config_wrap's concurrent-writers test, and the diagnosis the error name invites is wrong: the condition is contention, it clears in microseconds, and by the time the call is reached the bytes are already written and fsync'd. Aborting throws away completed work over a collision that resolves itself.

The retry is _filelock.replace_retrying, bounded at ten attempts 5ms apart, and it lives in _filelock rather than in the writer that happened to go red. That module already owns the primitives Windows spells differently — it exists because fcntl is POSIX-only and every call site used to drop its lock silently there — and os.replace is the second such primitive, not a config_wrap problem. Fixing it only where the gate failed would have left the gateway's per-tenant counters and the gateway key store, both of which have concurrent writers, broken the same way and unreported. All three atomic writers now route through the one helper.

A winerror outside the contention pair still raises on the first attempt, because retrying a genuine permission failure only makes an accurate answer slower. The POSIX branch is the bare call it always was.

The lock guarding a config depended on where distil was installed

The session lock exists to stop a release ("no live siblings remain, restore the original") from interleaving with a fresh claim ("register me, I'm writing my own config"). It only does that if every process wrapping one config agrees on one lock file. Keyed under DISTIL_HOME — which this changelog previously described as a tidiness fix — it did not: two distil wrap processes with different DISTIL_HOME values, from two installs or a test harness, took two different locks over the same config and serialised nothing. The race reopened silently, with the config the loser.

The anchor is a pure function of the config path again, beside the config as a sibling of the registry directory it guards. A shared per-machine directory keyed by a digest would also be pure, and is a worse trade on three counts: /tmp is swept and cleared on reboot while the registry and backup beside the config are not, so the lock could vanish while the state it protects survives; /tmp is world-writable on Linux, and a squatted lock path fails open(), which _filelock fails open on, degrading silently to no lock at all; and beside the config the kernel canonicalises for free, so two processes reaching one config through different symlinked parents contend on the same inode without a .resolve() having to be exactly right.

What this gives back is the leftover it was trying to remove: a <config>.distil-sessions.lock stays beside the user's agent config. It is accepted rather than cleaned, because it cannot be cleaned. Unlinking the file a waiter has already opened lets a third process create a fresh one and hold the "same" lock concurrently, so an unlink-when-the-registry-empties would trade a cosmetic problem for a correctness one. One empty file per wrapped config path, created once and reused forever, is the price of the lock working.

Every fix below is the same shape: a rule distil already enforces somewhere, not enforced on the surface that is actually exposed. The proxy refuses a body it cannot read; the gateway read it as empty. The client-supplied tenant header goes through a validator; the tenant claim from an identity provider did not. The proxy sets a socket timeout and says why in a comment; the gateway, the component documented as the one you bind to a network, set none. The key file and the temp file it is written through are created 0600; the blobs holding agent output were chmod'd 0600 a moment after. None of these were oversights of design — the design is written down and correct in each case. They were the second call site, and the second call site is where security bugs live.

A body the gateway could not read became the next request

The gateway sized every request body from Content-Length alone. parse_content_length returns 0 for a missing header, so a request framed with Transfer-Encoding: chunked was read as having no body at all — and its actual bytes stayed sitting in the socket buffer. On a keep-alive HTTP/1.1 connection the next parse picked those bytes up and treated them as a separate request, with its own separately-evaluated authorization, on a connection any front-end believed had carried exactly one. That is request smuggling, and it needs no exotic deployment: nginx, HAProxy, an ingress controller or a corporate egress proxy all honour Transfer-Encoding and will forward the chunked body faithfully. A reproduction sent one sendall and got two responses back, the second one a served dashboard.

The proxy has refused this since it learned to, in four lines. Those four lines are now httpguard.framing_rejection, called from both servers, and they refuse slightly more than before: any Transfer-Encoding rather than the literal string chunked, because the property that matters is that the body length does not come from Content-Length, and Content-Length and Transfer-Encoding together — the framing disagreement stated outright — as a 400.

Both servers also now close the connection on every rejection, which is the half of the fix a status code cannot do: the undrained body is still on the socket, and the only way it is never parsed is if nothing parses anything more. Framing was merely the loudest case. A cross-audit of this change pointed out that the oversized/malformed Content-Length 413 answers from the headers too — and so, it turns out, do the invalid-path 400 and every auth 401/403/429, all of which run before the body is read. So the close moved from the framing branch into _reject, the one function all of them already route through: one line, every caller, including the ones added later. The async proxy needs no equivalent change and did not get a cosmetic one — aiohttp parses framing itself rather than leaving the socket to the handler, and a direct attempt to desync it returned no response at all instead of parsing the trailing bytes as a request line.

A second pass found the rule was still not universal: the two rate-limit 429s in the gateway's auth path wrote themselves through _relay directly, so "every rejection closes" was true of every rejection that went through _reject and false of the ones that did not — on the path that by construction is being hit repeatedly. _reject now takes the extra headers a 429 needs (Retry-After), and all four 429s go through it. That is the actual lesson of this release stated once more: the rule is only as good as the number of call sites that can skip it.

A fourth pass found the last variant, and it needed no bypass at all — just the ordinary reading of a header. Both servers asked headers.get("Content-Length"), which returns the FIRST value and leaves the rest in get_all. So Content-Length: 5 followed by Content-Length: 0 was read as 5 here and may be read as 0 by a front-end that takes the last, and the five bytes stay queued as the head of the next request: CL.CL, the same desync reached without any exotic framing. framing_rejection now receives every Content-Length value and refuses the request when they disagree, including the 5, 0 comma list folded into one header line. Identical duplicates are still served — RFC 9112 §6.3 calls those one value sent twice, and refusing them would be a new outage wearing a fix's clothes. The async proxy again needed nothing: aiohttp's parser answers 400 to any repeated Content-Length, which a raw-socket probe confirmed rather than assumed.

Allowing those identical duplicates then produced a smaller bug of its own, worth naming because it is the same shape as everything else here: the guard said yes and the caller asked a different question. Both servers still sized the body from headers.get("Content-Length"), which for a repeat folded into one line hands back "42, 42" — a string int() refuses — so a perfectly well-framed request was answered 413 request body too large. The guard had already split those values apart in order to judge them, so it now returns the canonical one alongside its verdict and both servers parse that. One question, one answer, one place.

Transfer-Encoding was read the same wrong way, and the fix for Content-Length did not reach it: the guard still received headers.get("Transfer-Encoding"), the first value. So an empty Transfer-Encoding: ahead of a Transfer-Encoding: chunked read as "not TE-framed" while the body on the wire was chunked — the original desync, reached by adding one empty header. Both servers now pass every value, and a request is TE-framed if any coding is named anywhere, commas flattened so gzip, chunked counts. A present-but-empty header on its own still names no coding and is ignored, because 411-ing those would refuse requests framed exactly the way these servers require. aiohttp answers 400 to the duplicate form on its own and proxies the lone-empty one, so the async proxy again needed no change; both were confirmed with a raw-socket probe.

And the tenant validator, the one definition three doors were consolidated onto above, was anchored with ^…$ — where $ also matches immediately before a trailing newline. So acme\n was a valid tenant label, carrying into x-distil-tenant the exact character that splits a response header, which is the entire reason safe_tenant exists. The pattern is now \A…\Z. Putting the anchors in the pattern rather than asking four call sites to use fullmatch is deliberate: all four call .match() today, and a fifth written next year would otherwise inherit the bug for free. Two other ^…$ patterns applied to attacker-influenced input got the same treatment — Gemini path routing and sed-script classification in provenance. Of the three tenant doors only two were actually reachable: the header door already .strip()ed its value, and a header value cannot carry a raw newline over HTTP anyway. The OIDC claim and the operator key-issue path had nothing in front of them but the pattern.

Validating at issue() also only protects labels issued after the check exists. Every record already on disk — written by an older distil, restored from a backup, or edited by hand — walked straight past it on load and became the x-distil-tenant response header anyway. The key store now applies the same collapse to every tenant it parses, so the guarantee is about what is read rather than about when it was written. It warns once per process naming the file, and does not rewrite it: silently editing an operator's key file to make a warning go away destroys the evidence of how the label got there. The two other places a persisted tenant surfaces were checked and needed nothing — the dashboard escapes with html.escape and the Prometheus exposition escapes backslash, quote and newline. The response header was the one sink with no escaping of its own.

The chart's egress hole-punch list was half a list

The NetworkPolicy's default 443 rule is 0.0.0.0/0 minus a set of ranges, so anything missing from that set is somewhere a compromised pod can still send packets. It covered RFC1918, link-local and loopback, and stopped there. Now it also excludes 100.64.0.0/10 (carrier-grade NAT — EKS and GKE allocate pod and service CIDRs out of it, and some service meshes address sidecars there, so leaving it out left the cluster reachable), 0.0.0.0/8 (0.0.0.0 is a routable alias for localhost on Linux, i.e. a loopback bypass), 198.18.0.0/15 (benchmarking, used by Istio and some CNIs), 224.0.0.0/4 (multicast, where cluster discovery and gossip live) and 240.0.0.0/4 (reserved, and it contains the broadcast address). The list moved into values.yaml as networkPolicy.egressExcept with a line of prose per range, so the next person to read it can tell whether an entry is load- bearing; egressTo still overrides the whole rule for operators who can name their provider.

OIDC tokens have their own header now

Closing the OIDC gate above made a path reachable that had never carried real traffic, and it had a collision in it. The gateway read the OIDC token from Authorization: Bearer … and stripped that header before forwarding — correct for Anthropic, where x-api-key carries the provider credential separately, and broken for everyone else. For OpenAI, Azure and the Gemini bearer flavour that header IS the provider credential, and this gateway injects none of its own; it forwards the client's. So an OIDC-only deployment in front of OpenAI could not work at all: send the provider key and it fails JWT verification, send the JWT and the upstream receives no credential. The tests missed it because an echo upstream authenticates nothing.

The OIDC token now goes in x-distil-token, which collides with nothing and is stripped before forwarding like x-distil-key. Authorization: Bearer <jwt> still works where it always did — when the request also carries a provider credential, which is exactly the Anthropic shape — and is otherwise refused with a 401 naming the new header, rather than forwarded as a request certain to fail upstream. A bearer that is not a verified OIDC token is never consumed or stripped; it is the provider's. No upstream-credential injection was added: the gateway still forwards the caller's credential and holds none, which is the property that keeps it out of the blast radius of a compromise.

Configuring an identity provider did not turn on authentication

The gateway decided whether to require a credential before it decided what could serve as one. _auth_required() returned true when --require-keys was set or when a dsk- key had been issued, and the OIDC verifier ran inside that gate. An operator who pointed DISTIL_OIDC_ISSUER at their identity provider, set a signing secret, issued no gateway key and did not pass --require-keys was running a completely open gateway while every OIDC setting read back as configured — the one failure mode where the configuration surface actively tells you the opposite of the truth. A configured issuer now makes authentication required on its own, read from the environment on the same call the verifier uses so the two can never disagree about whether OIDC is on. Closing the gate must not close the door: with an issuer configured and no key store at all, a verified bearer token is now accepted as a complete credential rather than rejected for the absence of a store that deployment does not need.

A tenant label is a header value, and header values have no escaping

The tenant a request is booked under reaches three places that assume it is inert: an x-distil-tenant response header, the dashboard, and the per-tenant accounting map. BaseHTTPRequestHandler.send_header performs no CRLF validation. The client-supplied header path has always been checked against a bounded, punctuation-free pattern; the OIDC claim path was not checked at all, so a tenant claim of acme\r\nX-Injected: yes was emitted verbatim into the response. It takes a validly-signed token to reach, so the attacker already holds a credential — but where the identity provider lets a user influence that claim and a shared cache sits in front, splitting a response is worth more than the token. The pattern now lives in authz.TENANT_RE, with one definition for all three doors a tenant label comes through. An unsafe claim collapses to a stable digest rather than falling back to sub, which comes out of the same token and would sanitise nothing.

The third door was distil gateway keys issue --tenant "". The empty string is the exact sentinel the auth path uses to mean "401 already sent, stop", so a key issued to that tenant authenticated correctly and then every request it made returned no response at all — not an error, nothing. Issue-time validation refuses it, along with every other label the response header could not carry.

The exposed server was the one without a socket timeout

proxy.py sets a 600-second client timeout on the handler class, with a comment explaining that StreamRequestHandler.setup() applies it to the accepted socket so no read or write to a client can block forever. The gateway handler set protocol_version and nothing else, which leaves the timeout at None: a peer that opens a connection and stops sending, or declares a large Content-Length and dribbles, holds a threading-server thread for the life of the process. The gateway is the component meant to be bound to an interface. It now carries the same value from the same constant.

Owner-only means at creation, not a moment later

atrest.py creates the master key by passing the mode to os.open through an opener, and gateway_keys.py does the same for the temp file it writes key hashes through; both carry comments recording the measured 0644 window that motivated it. The MCP handle store and the restore blobs — the files that hold the actual original tool output — were still write_bytes followed by chmod, which leaves them at the process umask for the duration of the write. A poller caught 0644 on a restore blob during a 400-write loop. The contents are ciphertext, so ordinarily that window leaks nothing readable; the sharp edge is DISTIL_NO_ENCRYPT_AT_REST=1, which the threat model recommends for ephemeral homes and which makes those same bytes plaintext. The opener is now atrest.write_owner_only, shared rather than copied, and both MCP write sites use it. The receipt chain was the last file of this shape, and it appends rather than rewrites, so it takes the same mode through atrest.owner_only directly. That closes the class: every store distil creates now gets its mode from the os.open that creates it, and none of them depends on a chmod arriving in time.

Controls the CI did not have

live-cert.yml holds ANTHROPIC_API_KEY and used three third-party actions by mutable tag while release.yml — the workflow that publishes — has been fully SHA-pinned all along. The same three actions at the same versions are already pinned there, so the SHAs were copied rather than looked up. A pin without a bump is a pin to something old, so dependabot.yml now watches GitHub Actions and pip weekly, which is what rewrites a SHA and its version comment together.

A new supply-chain.yml adds four checks that answer questions no existing gate asks: a pip-audit of the resolved optional extras (the core declares no runtime dependencies, so auditing the project itself would pass forever by auditing nothing), gitleaks with an in-repo .gitleaks.toml so a local run and CI agree on the three known false positives, CodeQL for Python and JavaScript, and npm audit for the two packaged Node surfaces. Both npm packages are dependency-free today and neither has a committed lockfile, so that job resolves one first — it is a tripwire for the day a dependency is added, not a claim that anything is being audited now. It runs weekly as well as on pull requests, because a CVE is published against code that has not changed and a gate that only runs on a diff can never see one.

A NetworkPolicy that permitted every destination

The Helm chart's egress rule for HTTPS listed a port and no to: selector, which in NetworkPolicy semantics means every destination on that port — while the comment directly above it said that a compression proxy able to reach arbitrary hosts is an exfiltration path. The default is now the narrowest selector that works without knowing the operator's provider: the public internet minus RFC1918, link-local and loopback, so a compromised pod cannot reach the cluster, the node, or a cloud metadata service on 443. values.yaml carries an egressTo override and says plainly that narrowing it to your provider is the point.

Added

  • Mistral Vibe is wrapped — distil wrap -- vibe. Vibe sat on the "could not verify" list because its endpoint lives in a config.toml whose shape no docs page publishes. Patching that file was never the answer. Vibe's own ADR 0005 puts a VIBE_* environment layer above both the user and project TOML layers, and its config layer reads that layer with env_prefix="VIBE_" — so the schema's providers list is VIBE_PROVIDERS, taking a JSON array, and entries merge across layers on name. wrap exports a one-element array redirecting the mistral provider and leaves the rest of your configuration alone. Nothing on disk is touched, so there is nothing to restore and nothing a crash can leave behind. It is the first preset whose variable holds a document rather than a URL: AGENT_ENV_TEMPLATES renders $BASE into a literal, because exporting a bare URL where the agent expects JSON is the exact failure this project refuses to ship — wrap would report success, start a proxy, and route zero traffic.
  • The Cline CLI is wrapped — distil wrap -- cline. It was declined for having no published config schema. It has one; it is just not on the docs site — the zod StoredProviderSettings in cline/cline, with a committed fixture of the real file. ~/.cline/data/settings/providers.json, where providers.<id>.settings.baseUrl is documented in the code as outranking both the regional API line and the provider default. The preset also sets lastUsedProvider, because an entry the CLI never selects routes nothing — and it honours CLINE_DATA_DIR, because patching a file your CLI does not read is the same lie by a different route.
  • The Kilo Code CLI is wrapped — distil wrap -- kilo — via KILO_CONFIG_CONTENT, not a config file. This began as a config-file preset patching ~/.config/kilo/kilo.json, and that was wrong for a reason worth recording. Kilo's own precedence table puts global config files at 4 and a project-local ./kilo.json at 6, so inside any repo shipping its own config the patch landed on a file the child never read, while wrap reported success. That is the exact failure this area exists to prevent, reintroduced by a fix for it. The same table lists KILO_CONFIG_CONTENT at 8, above both, and Kilo's loader hands it straight to loadConfig(text, …) as config content. So the preset exports a config document instead: nothing read, written, backed up or restored, no project file able to shadow it, and a kilo.jsonc full of comments never at risk of being rewritten without them. What goes in that variable is a base URL for Kilo's built-in anthropic and openai providers and nothing else — it declares no models of its own, because Kilo treats a custom model with no limit.context/limit.output as having limits of zero, and a provider that is selectable and then quietly mismanages context for a whole session is the same half-working shape as a guessed variable. Overriding the built-in ids keeps Kilo's own catalogue, real limits included, and changes only the endpoint: nothing to pick by hand, your top-level model untouched, and no credential in the environment since the key variable comes from the catalogue too. A session on some other provider (OpenRouter, Kilo's own, a local gateway) is simply not redirected — those are not wire shapes distil speaks.
  • distil wrap --list (and --json). Every target, its mechanism (environment variable / config file), the provider wire shape distil has to speak for it, the routing knob, and the primary doc that contract was read from with the date. The agents wrap cannot reach are on the same list with the reason — that half is the useful half.

Fixed

  • A preset can export the right variable and still route nothing, because the variable's own client uses it a way distil never checked. Kilo Code's KILO_CONFIG_CONTENT named the right env var but the wrong value: its provider layer forks @ai-sdk/anthropic / @ai-sdk/openai, both of which use a configured baseURL literally and append only the leaf path (/messages, /chat/completions) — so $BASE alone landed every request on /messages, a path is_compressible_path does not recognise, while wrap reported success. The fix appends /v1 in the template, not the base preset. The same defect class was then checked against every other AGENT_ENV_TEMPLATES entry and confirmed live (real installs of openai-python, openai-node, and litellm, none of which insert /v1 for an explicitly-set base_url either) against aider, OpenCode, and Qwen Code — all three built on that same literal-base_url convention, all three now exporting $BASE/v1. tests/test_reach_contract.py pins the fix per SDK convention (with its own doc citation per row) by running each preset's exported value through the real proxy against a fake upstream and asserting the request both lands on a path the proxy compresses (b) reaches the upstream on that same path, and (c) actually triggered the compression branch — a genuinely-compressible tool result in the canned body must come back with x-distil-tokens-saved > 0, not just a passthrough that happens to land on a compressible-shaped path. Reverting any template's /v1 fails it.
  • grok, kimi, and openhands had the exact same defect, confirmed this round from each client's own source (not guessed at): xai-org/grok-build's resolve_inference_base_url() and MoonshotAI/kimi-code's packages/kosong/src/providers/kimi.ts both use their base URL literally, and both default it to a value that already carries /v1 — so a distil upstream default of .../v1 plus a bare $BASE export would double the segment into .../v1/v1/..., a 404 that reads like a distil bug, once the OpenAI-SDK-style /v1-autoinsert assumption underneath the old preset stopped holding. AGENT_ENV_TEMPLATES now exports $BASE/v1 for both, and AGENT_PRESETS's upstream is stripped back to the bare host so the proxy's own forward doesn't double it either. OpenHands turned out to be the same convention one layer down: LLM_BASE_URL is forwarded into LiteLLM's api_base verbatim (confirmed from OpenHands/software-agent-sdk's own docstring: the resolved value LiteLLM would otherwise compute is deliberately discarded so a later per-call resolution isn't frozen), and LiteLLM injects no fallback base for its openai provider branch — so OpenHands now gets the same $BASE/v1 template as aider.
  • Codex's and OpenCode's OpenAI presets were modelled on the wrong wire shape. Both are Responses API, not Chat Completions, confirmed from two independent sources: codex-rs removed wire_api="chat" entirely (codex-rs/model-provider-info/src/lib.rs), and @ai-sdk/openai@4.0.75's bare openai(modelId) invocation (no .chat/.responses suffix, which is how OpenCode's own packages/llm/src/providers/openai.ts calls it, by independent default) now resolves to createResponsesModel. AGENT_META's shape for both is corrected; tests/test_reach_contract.py gained an openai_responses case and body shape (function_call_output items) to prove it.
  • A genuinely deeper finding on codex, surfaced while chasing the shape question and left unfixed pending an answer, per this project's own rule against guessing: codex-rs is a native Rust client, not the openai-python/-node SDK the old preset comment assumed, and it builds request URLs by literal concatenation (Provider::url_for_path). Its base_url field, though, is populated only from the TOML openai_base_url config key (codex-rs/core/config.schema.json) — no env-var-to-config-field mapping for OPENAI_BASE_URL was found anywhere in codex-rs/config. Distil's codex preset may not route codex's traffic at all, independent of any /v1 question. Left unchanged; documented in AGENT_META["codex"]'s note and this file rather than silently patched.
  • Unverified and deliberately unchanged (no live-confirmed answer for what the client does with a bare base_url, so left as-is rather than guessed at): goose, copilot. OpenCode's real end-user override is a config file (opencode.json), not a plain env var, and whether OPENAI_BASE_URL genuinely outranks an explicit config-set baseURL (versus only supplying a fallback when none is configured) was not re-verified this round — flagged in AGENT_META["opencode"]'s note for follow-up.
  • A tautological regression test. test_kilo_fix_is_load_bearing monkeypatched AGENT_ENV_TEMPLATES["kilo"] to a locally-defined reverted string and then read that same entry back out — it never read the real template and never sent a request through the proxy, so it could not have caught the regression it claimed to guard. Deleted; the kilo-anthropic/kilo-openai parametrized cases in the table above are the real guard.
  • /v1/v1 documented, not patched. httpguard's _CHAT_RE/_RESPONSES_RE are anchored, and the proxy forwards _upstream + path unchanged, so a client whose base_url already carries /v1 and appends its own /v1 leaf on top lands on a doubled prefix neither regex matches — an uncompressed passthrough today, not a silent drop and not a match on a malformed path. tests/test_reach_contract.py now pins that behaviour directly rather than leaving it implicit; the allowlist itself is unchanged.

Changed

  • Config-file presets follow the path the child was actually told to use. A preset now receives the wrapped command's argv, because several of these tools let you move their config, and patching the default then writes a real file nobody reads — wrap reporting success while routing nothing, the same shape as the Kilo shadowing bug. The Cline CLI publishes three such knobs at three different depths: --config (the settings directory itself), --data-dir (two levels above it) and CLINE_DATA_DIR (one level above it). Each is honoured on its own. When more than one is given, Cline's reference documents no precedence between them, so distil names them and refuses rather than guessing. Crush follows XDG_CONFIG_HOME for the same reason. The Continue CLI refuses the mirror-image case: its preset injects by appending its own --config, so a user who passed one would be silently overridden, or silently lose to it. Crash recovery sweeps the flagged path too, so re-running the same command cleans up after a kill -9 under those flags. Factory Droid and Oh My Pi publish no relocation knob, so there was nothing to follow.
  • A second distil wrap of the same config-file agent is refused, not silently fought over. The per-session registry made the shared backup safe — whose bytes to restore and which session restores them — and that had been mistaken for making concurrency safe. It never was. Every config-file preset writes the same active provider entry pointing at its own proxy port, and one file cannot name two ports. Two live wraps of, say, Crush gave: the second repointed the file at its proxy, so the first agent's traffic ran through the second session and landed in its ledger; the first exited, its proxy died, and the config the second was still using named a dead port; the second exited last and restored the pre-wrap backup under a session already gone. Having the second reuse the first's proxy cannot fix it either — the first owns that proxy's lifetime and takes it down when its agent exits. distil wrap now checks for a live holder before starting a proxy or writing a byte, exits non-zero, names the pid holding the file, and points at distil default --always-on, which is how several agents share one long-lived proxy with no config patching at all. The check is repeated inside the same lock as the claim, so two wraps starting in the same instant cannot both pass it. Dead registrants are reaped exactly as before: a kill -9ed session never blocks the next wrap.
  • One catalogue, generated docs. The same three facts about each agent were restated in five places — two preset registries, a dict in cli.py, and prose tables in README.md, docs/IDE-AGENTS.md and docs/integrations.html — and they drifted. distil/targets.py now joins the registries with their doc metadata and scripts/build_agent_tables.py renders all three documents from it; tests/test_wrap_targets.py fails when any of them goes stale. A preset added without its cited source is now a test failure rather than an undocumented target.
  • Stale claims about other people's tools, corrected. The docs said Warp had "no published base-URL override at all"; Warp ships a custom inference endpoint now. The real reason it is out of reach is the better one: the agent harness runs on Warp's own servers and its docs reject localhost and private addresses, so a local proxy can never be the target. The docs also called Cline "not a CLI" after Cline shipped one — that CLI is wrappable as a process, it just publishes no base-URL knob it honours. And httpguard.py described Azure OpenAI's /openai/v1/ surface as preview with a required api-version; it is GA now and the parameter is optional. The path patterns were already right either way, since the query string is stripped before they run.
  • Copilot's BYOK contract, re-checked 2026-09-16. COPILOT_PROVIDER_BASE_URL / _TYPE / _API_KEY are unchanged, so the preset stands. Two things it deliberately does not set are now documented instead of silent: COPILOT_MODEL is required and only you know which model you want, and Azure OpenAI needs _AZURE_API_VERSION plus _WIRE_MODEL (your deployment name) alongside the resource-shaped base URL.

Not added, on purpose

Seventeen targets were checked against their own primary documentation. Fourteen got no preset, and docs/IDE-AGENTS.md now lists every one with the page and date that was read. Three distinct reasons. No knob at all: Cursor CLI, Amp, auggie, Antigravity and Tabnine publish no base-URL override — their only network setting is a whole-process HTTP_PROXY. A knob that cannot be local: Warp, above; and Cody, whose override is an admin setting on the Sourcegraph instance rather than on the client. A real knob, but nothing that scopes to one session: Roo Code keeps its profiles in VS Code's Secret Storage rather than a file; OpenClaw's baseUrl belongs to a Gateway daemon its own README describes the CLI as merely connecting to, shared with chat channels and, on a team install, other people; ZCode is a desktop app with no process to launch; and Zed's built-in Anthropic provider documents no api_url at all, so redirecting it means adding a second provider the user must pick by hand, in a file Zed's own settings page rewrites while it runs. Trae and Junie could not be verified at all: every docs path returns the same client-rendered shell over a plain fetch. Bedrock's SigV4 path is out of scope.

Note what is not a reason: "the settings file is global." That was the stated ground for declining Cline and Kilo Code, and it was wrong — a shared config file is precisely what config_wrap claims, backs up, patches and restores for Crush, Oh My Pi and Factory Droid already. Both are presets now. What is disqualifying is narrower and sharper: a global file that some other file outranks for the directory the wrap runs in, which is what sent Kilo to an environment variable rather than to this list. What remains declined is declined on the specific mechanism, because a preset built on a guess is indistinguishable from one that works right up until you check the savings counter.

The ninth command still knows the other eight exist

Docs and CLI-surface polish; no runtime behavior change. The through-line: in every case here the accurate answer already existed somewhere in the codebase, but only if you already knew where to look — --help for one savings command never mentioning its siblings, a CLI reference example nobody had actually run through argparse, TOST printed with alpha/margin and no word on what either means, onboard--no-interactive and offboard --no-interactive inventing their own wording for the same rule.

  • distil stats, dashboard, dissect, doctor, shadow-stats, and receipts now cross-reference each other in --help, and onboard's closing summary and every session's proof-ledger print end by naming the exact next command (distil stats / distil dissect <session>) instead of a vague "check your savings."
  • onboard --no-interactive and offboard --no-interactive now state the identical --yes-vs-no-interactive rule verbatim, so learning it on one command means already knowing it on the other; both gained a Flags table in the CLI reference.
  • certify/bench's --margin/--alpha flags, and the TOST and Tier-0/1 lines they print, now say in plain words what they mean rather than assuming the reader already knows two one-sided tests; dissect's terminal glossary now defines decision-equivalence and prefix replay, both used earlier in the same report.
  • docs/cli.html's own $ distil … examples are now parsed by the same test that checks every other doc's — it caught two that never actually ran (certify-provider was missing its required episodes argument and had a --runs flag that doesn't exist; receipts --export took a value it doesn't accept). Both fixed.
  • distil dissect --serve and distil dashboard --web now document their CI guard: under CI they print the URL and exit rather than block, and --foreground overrides it. This used to key off "no TTY", which also caught nohup, a systemd/supervisor unit, and IDE run tasks — legitimate non-interactive launches that do want the server; it now checks for CI by name (CI, GITHUB_ACTIONS, GITLAB_CI, BUILDKITE, TF_BUILD) instead.
  • README gained a Proof and provenance section (PEP 740 attestations, the CycloneDX SBOM, the weekly OpenSSF Scorecard run, the threat model and security whitepaper, distil validate) and a link to docs/llms.txt for agents reading the repo instead of a human.
  • LangGraph has its own integrations page (docs/langgraph.html) alongside LangChain's, and the Vercel AI SDK page now documents distilMiddleware() next to the proxy route it already covered.
  • The Helm chart's default port matched distil proxy, not the distil gateway it actually deploys — fixed in values.yaml, chart version bumped to 0.1.2.

Fixed — numbers the artifacts did not support

The site said ✓eq 99.5% in the hero terminal, under a caption reading "Real, reproducible output." It was not output. distil/cli.py prints that check glyph only at eq >= 0.99, and only past the 50 A/B + 30 A/A reporting floor, and the maintainer's own traffic had never cleared it. The same invented verdict appeared a second time in the trust card on the same page, and twice more in plugins/distil/README.md with a sample size (1.2k) about three times the real one. All four now show the reading the estimator actually produced on 2026-09-15: 97.5% with a 95% CI of [95.5, 99.5] over n=398 A/B and 399 A/A, paired difference −0.025 [−0.045, −0.005], digest and lossless-only mixed. That is under 99%, so every surface shows the warning glyph, not a check. Replays run hot — 399 of 399, temperature not pinned — so the paired difference is the statistic and raw agreement reads 53.0%; the artifact says so in its own header. The sample now clears the floor, so the four places that said "the current live sample is below that floor" say what it is instead. Artifact: benchmarks/results/shadow-live-2026-09-15.json. The 2026-09-04 reading (44 A/B, below the floor, unpaired estimator) stays committed as the evidence for the 1.13.0 withdrawal.

docs/benchmark.html's certified-frontier block read 52.3% for the lossless rung. The log it is presented from, docs/paper/results/derc_live_compare.2026-07-05.log, says 47.9%. Every other row in that block matched the log exactly, which is why nothing noticed a single drifted digit pair inside an otherwise faithful transcript. Fixed to 47.9%.

The July head-to-head — 83.2% savings at 0% decision-change, LLMLingua-2 53.1% flipping 1-in-8, Headroom 39.7% — is real and reproduces from its log, but it was the undated headline on nine surfaces while the competitor it names had shipped ten minor versions. It is now dated inline everywhere it appears: 2026-07-05, distil 1.10.1 vs llmlingua 0.2.2 and headroom-ai 0.27.0. Beside it, the 2026-09-04 re-run against headroom-ai 0.37.0: distil-causal 52.9% tokens / 58.7% dollars at 100% decision-equivalence (PASS) against Headroom's 1.7% / 2.0% / 81% (FAIL) on the warm corpus gate, and 35.6% tokens for +4.9% dollars on the read→edit→re-read codebench workload. The "2.1× less aggressive" ratio was removed from every page rather than re-dated: it is 83.2/39.7 from the old run, and it does not survive the re-run in either direction. The ledger no longer marks the identical claim verified on one page and stale on two others.

The adoption page's trust ring printed a capped value as a measurement. distil/census.py sends round(min(100, pct)) because the collector rejects anything above 100, and the paired estimate genuinely can exceed 100% when compression agrees more often than the model agrees with itself. The cap cannot hide harm — the difference is bounded below, so only the upper bound can bind — but an unlabelled 100% is the one number shape this project exists to criticise. The ring now prints 100% (capped) and says why in the caption beneath it. The label is a word rather than a ≥, because a rate printed above 100% reads as broken to anyone not holding the paired estimator in their head. The wire format is unchanged.

Fixed — a claim whose artifact directory was empty

docs/claims.json cited benchmarks/results/2026-09-06/ for the re-read delta's 31.4%→41.7% headline. That directory held one README and no artifact: benchmarks/.gitignore ignores *.out, the 2026-09-04 batch was force-added, and this one was not. The run was repeated on 2026-09-15 with the README's exact commands and the outputs are force-added now.

The headline pair reproduced exactly: 31.4% → 41.7% tokens and 35.8% → 49.1% dollars under the client shape that bills. Two secondary figures did not, because "after" is now main at 1.53.0rc1 rather than the feat/reread-delta branch, and the site takes the re-run: the unmarked client shape reads +12.2% dollars, not −15.4%, because forwarded-bytes prefix replay stops charging it for a prefix it never cached; and distil validate reads 175/175 over 25 cases, not 150/150. distil bench is byte-identical before and after, as the README claimed. The results README names both superseded values rather than overwriting them.

Changed — the claims gate reads artifacts and scans pages

docs/CLAIMS.md has always stated the rule in bold — no number on the site without an entry in docs/claims.json — and nothing enforced it. The gate checked that ledger entries still matched pages; it never checked that pages were covered by the ledger, and it never opened an artifact. Three failure states were live at once and all three passed CI: an artifact path that did not exist, an artifact that contradicted the page, and the site's most-repeated claim carrying no artifact key at all.

tests/test_claims_coverage.py closes both directions. It scans README.md, docs/*.html, docs/llms.txt and plugins/**/*.md for percentages and multipliers and fails on any that no entry naming that page mentions; sample terminal blocks, fenced code and statistical notation are stripped, and the handful of genuine non-claims (a provider's published cache-price ratio, a status-line threshold, someone else's marketing quoted in order to refuse it) each carry a one-line reason. It also opens every artifact: the path must exist, and where the entry lists values those values must appear in it, with JSON fields compared as both fractions and percentages so a page's 36.8% matches a stored 0.368. Entries that genuinely cannot be machine-checked — a ratio of two columns, a bound computed at report time, a command whose output was never committed — are marked "check": "manual" and must say why.

Running it turned up the rest of this entry's work: the ledger grew from 37 entries to 50, llms.txt and plugins/ went from having no coverage at all to being fully scanned, and the pages that carry the product's load-bearing claims — index.html, benchmark.html, benchmarks.html, output.html, architecture.html, concepts.html, faq.html, getting-started.html, llms.txt, README.md, plugins/** — are clear. Pages not cleared in this pass carry a frozen debt list that may only shrink: a new number on them still fails CI. The ledger's own reproduction note for the adversarial battery had drifted too (28 cases / 175 checks against a real 32 / 231) and is corrected; note fields are still not machine-checked, which is the next thing to fix.

The next stale number trips CI instead of a reader.

Fixed — a dashboard card quoting a floor the code had retired

The HTML ledger's decision-equivalence card told a reader below the floor that it "needs 25+" shadow samples. There is no 25 in the code: the shared reporting floor is 50 A/B plus 30 A/A, and the card is only ever reached with a rate that already cleared it. The card now reads those two constants from distil.shadow, so it names the floor it actually enforces, and the three tests that pinned the retired wording assert against the constants rather than a copied string.

The same retired floor was still quoted in five more places, all of them found by grepping for it rather than by any gate: the terminal dashboard's collecting line told the reader to "need 25" while the card one screen over had been corrected, render_html's docstring and the status-line comment in distil/cli.py both named 25 as the threshold, two status-line examples rendered a de 12/25 that the code can no longer produce, and docs/BETA.md asked beta users for ≥25 samples before reporting. All now read the floor from VERDICT_MIN_AB/VERDICT_MIN_AA or spell it out as 50 A/B + 30 A/A.

Fixed — the dashboard could publish a rate with no noise baseline behind it

Every surface gates decision-equivalence on the paired verdict: VERDICT_MIN_AB A/B samples and VERDICT_MIN_AA A/A ones, because a rate with no self-agreement baseline cannot tell compression harm from the model's own nondeterminism. The terminal dashboard did not. It took a bare change_rate and an A/B count, gated on the count alone, and was fed ShadowLedger.rate() — the raw A/B rate. A ledger with 50 A/B rows and zero A/A rows printed a confident percentage that the status line and shadow-stats, reading the same ledger, both refused to state.

Both renderers now take the Equivalence object itself rather than a rate and a count, so the threshold lives in one place and a page cannot hold an opinion the verdict does not. Below the floor they print which counter is still short, from a new Equivalence.shortfall shared with any future surface. It names the constraint that binds: once both arms are full a paired verdict can still be short on the paired pool itself, and showing "50/50 A/B, 30/30 A/A" beside a refusal reads as a bug rather than as evidence still accruing.

shadow-stats --json keeps its raw decision_change_rate, which is correct — it sits beside equivalence_pct, below_reporting_floor and the paired fields, each labelled for what it is.

Two surfaces also read the ledger unscoped. A verdict belongs to the signature algorithm that produced it — SIG_VERSION is bumped precisely so old rows are never compared against new ones — and the status line, leaderboard, census feed, proof ledger and web dashboard all pass current_only=True. The terminal dashboard and distil doctor's shadow check did not, so a ledger holding only rows from a retired algorithm could carry either past the floor and print a confident percentage that every other surface, reading the same file, declined to state. Both now scope. shadow-stats still defaults to the current signature and widens only under --all, which is the one place reading retired rows is the point.

The doctor check's own tests had stubbed ShadowLedger.load with a lambda taking no arguments, which is part of why this went unseen: the stub could not have noticed the missing scope. The new tests for both surfaces write real rows to a real ledger and read them back through the real command.

Fixed — the claims gate skipped the whole plugins tree on Windows

Page keys were built with Path.relative_to, which renders with the host separator. On the Windows runner the plugin pages arrived as plugins\distil\README.md, matched no ledger entry spelled with forward slashes, and the new coverage gate reported every number on them as uncovered. Keys are now normalised to posix on both sides — the scan list and the page fields read out of docs/claims.json — so the gate compares the same spelling everywhere, with a unit test that feeds it a backslash path.

Claude Code's connector search was switched off by distil

Claude Code defers MCP tool definitions by default and loads one only when the model searches for it. It turns that off whenever ANTHROPIC_BASE_URL names a non-first-party host, "since most proxies don't forward tool_reference blocks". Every distil wrap -- claude is exactly that, so every connector's full schema rode along on every turn. On the maintainer's machine that meant a tools array of roughly 390k tokens per request, nearly all of it one connector the agent called 5 times in the nine days measured.

  • distil wrap -- claude now sets ENABLE_TOOL_SEARCH=true for the child, through setdefault, so an exported ENABLE_TOOL_SEARCH=false still wins. distil forwards defer_loading, tool_reference blocks and the beta header byte-identical. A new test pins all three with the digest path active. This is Claude Code's own first-party default, not a distil transform, so it applies on subscription traffic as well.
  • The request record no longer counts a defer_loading: true definition as overhead (it is not billed) and records how many there were (tools_deferred). It also records, by name only, the definition tokens per MCP server (mcp_servers) and which servers the conversation has called (mcp_called).
  • distil discover gains unused_connectors: MCP servers whose definitions were sent on 20 or more requests and never called anywhere in the window. Each one is priced where it sits in the prompt (cache read first), not at the flat input rate.
  • An mcp-compressor-style "hide unused tools behind a meta-tool" was modelled on real traffic and not built. distil is not the MCP server, so every unlock would rewrite the cached prefix, and the native mechanism removes the same tokens without that cost. Numbers, and the conditions for reopening, are in docs/adr/0013-unused-connectors-are-claude-codes-to-defer.md (benchmarks/results/2026-09-24/lazy_tools_model.json).
  • distil default --always-on writes ENABLE_TOOL_SEARCH=true into the same Claude Code settings env block as its ANTHROPIC_BASE_URL pin, but only if the key is absent: a value you set (true, false, auto:N) is never touched. distil records the files it added the key to (settings-added.json in the distil home). --undo, distil offboard and the uninstall.sh escape hatch remove the key only from those files, and only while it still reads true. If you have changed it since, it is yours and stays. If you delete it, the next distil run that sees it gone records that, and distil never adds it back.
  • Ownership is recorded only after the settings write succeeds, so a failed write can never make undo delete a key distil did not write. The pin and the key go in as one read-modify-write, and --undo takes both out the same way: one write and one .bak of the original per file. The pin is removed first and independently, so trouble with the tool-search key can never leave behind the base URL that kills sessions.
  • Every Claude Code settings write distil makes, and every rc-file write, is now atomic. This includes the existing ANTHROPIC_BASE_URL and status-line paths.
    • The temporary file is created 0600 (settings can hold API keys), fsynced, then given the target's mode and, best-effort, its owner before the rename.
    • It is removed if anything fails, so no copy of the file is left behind.
    • The write goes through a symlink to its target.
    • A brand-new file is created 0600.
    • A malformed settings file, or a symlink loop, is reported and left untouched.
  • Known limit: if you delete distil's key and re-add the identical "true" by hand with no distil command run in between, the two are indistinguishable, and undo removes it.
  • Downgrading distil below this release leaves ENABLE_TOOL_SEARCH set. That is harmless against Anthropic's API, since it is Claude Code's own default there. Behind a gateway that strips tool_reference blocks, it breaks loading of MCP tools. To remove it by hand, delete the "ENABLE_TOOL_SEARCH": "true" line from the env block of your Claude Code user settings. Or run distil default --always-on --undo before downgrading.
  • Verified live (2026-09-25). A wrapped Claude Code session that called an MCP tool recorded tools_deferred of 4–5 on every request (tool payload 9,733 tokens, versus ~390k-token arrays before), prompt-cache reads intact, no request failures. If it breaks, there are two ways to revert:
    • Per user: export ENABLE_TOOL_SEARCH=false (wrap) or set it to false in the settings env block (always-on). Both win over distil's default.
    • In code: AGENT_PRESETS["claude"] back to {}, and cmd_default back to wire_settings_env for the pin instead of wire_always_on_settings.

distil discover — the report that says what to do next

Every ingredient for "here is where you are still leaving savings on the table" was already being computed, per session, by distil dissect: the fixed-overhead share and the per-tool cost of it, which sessions never reached the digest tier, how often the cached prefix drifted and what the provider re-billed for it, how many tokens were re-folded after first sight, how far the system prompt grew. What was missing is the only part a user actually needs — the cross-session aggregation, ranked, with a number and a command against each line. discover adds no instrumentation and no estimator; it reads the same content-free records through the same accessors and does arithmetic on them.

Two constraints shaped it more than the detectors did. The first is that a best case must never be readable as a typical one — the audit of this field found the gap between a competitor's headline and its fleet median to be the recurring trust problem, and a savings tool that can be read that way has no standing to point it out. So the median and the p10/p90 of per-session savings print beside the best session, labelled. The second is that an estimate may not invent its own ratio. Where a line needs "what would the other mode have been worth", it uses the rate this machine measured, in tiers: its own recent sessions first, then its lifetime history, and only a machine that has never run that mode at all reaches the published benchmark figure — and then the line names the corpus and says it is not your traffic. Detectors that cannot measure return nothing rather than something: churn is counted only over sessions whose cache-read share was measured and low, because a resend the provider already discounts is not a saving, and an unmeasured one is neither cheap nor expensive; a drift ratio needs four comparable turns before it is a ratio at all. "Measured" means the provider's own usage object carried the cache-split field — a literal 0 it reported is a real zero, but the field's outright absence (true of every OpenAI/Gemini row, and of every row written before this release) stays unmeasured rather than being read as one; the two are opposite diagnoses and conflating them would have inflated churn for exactly the providers that never send the split at all. A dollar figure is refused the same way when most of a window's tokens came from a model the proxy could not price at all — an unpriced minority may not dilute a priced majority's rate, and an unpriced majority gets no dollar figure rather than one extrapolated from an unrepresentative sample. "Nothing to recommend" is therefore a result, not a failure to look, and it is only printed when a detector actually ran: a window of ledger-only or unbooked sessions says so instead of a false all-clear.

The wrap-exit proof ledger gains at most one line, and only when a detector actually fired — an advisory that prints "0 actions" on every exit is an ad, not a finding. dissect() grew two keyword arguments (ledger_rows, shadow) so that dissecting twenty sessions stops re-parsing a 37,000-run ledger twenty times over, and a third (since_ts) so that --since N bounds an always-on session's rows to the window asked for rather than folding its whole history in; nothing else about a dissection changes, and there is still one implementation of it.

Every wrap now ends with a verdict that can come back negative

Every piece of statistical machinery in this repo already worked. None of it was ever shown to the person whose traffic it was measuring. The e-process in drift.py had no caller. The conformal bound was reachable only from a command nobody runs twice. The receipt chain was hash-linked and verified on request, which means in practice never. The gap was not rigor; it was that rigor stayed in the library and the user got a savings number.

Every distil wrap session now ends with four verdicts, and the same four print in distil stats and distil dissect from one shared function, so three surfaces reading one ledger cannot disagree about it. Each one can come back negative — that is the entire reason to print it.

  • budget — the certified decision-change budget, checked after every request. Live shadow rows feed a persisted betting e-process (~/.distil/drift.json, folded once per row in file order, locked). Capital crossing 1/δ means the live decision-change rate has exceeded the certified 5% at 95% confidence, and the line says BREACHED at samplek until the evidence is archived by distil reset --shadow. Ville's inequality is what makes peeking free. The state is bound to the stream it counted — a fingerprint of the rows already folded, plus the signature version they were read under — so archiving or truncating shadow.jsonl outside distil reset --shadow, or a SIG_VERSION bump that filters old rows out, rebuilds the e-process over the current stream and says restarted: the shadow stream was replaced, instead of quoting a stale n while ignoring every new sample until the file outgrows the old count. The gate is a build gate, not a claim: 2,000 null runs at exactly the budget alarm 99 times (4.95% against a δ of 5%), and a true rate ten points over budget is caught on 500 of 500 runs within a median of 172 requests.
  • risk — a distribution-free upper bound on the same losses. Deliberately wider than the bootstrap interval beside it: the bootstrap estimates where the rate is, this one states where it is not, assuming no distribution and holding at finite n. A 1,000-run coverage simulation gates it.
  • output — what compression did to reply length. Shadow already measured it and only dissect showed it. A shorter prompt that buys a longer answer can cost more than it saved. The direction word prints only when the 95% interval excludes zero; it is the measured effect on ordinary traffic, not --shape-output, which asks for shorter replies.
  • receipts — the chain, verified at exit. distil receipts --verify now exists as the obvious spelling of the default action. Verifying a broken chain used to report receipt 1 of 2 when three receipts existed, because the scan stopped counting at the break; the second number is the one that tells a reader how much of the artifact is in question, so it now counts the file. Writing a receipt is a read-modify-write, so the head read and the append are now one locked critical section — unserialized, two concurrent requests both claim the same predecessor and fork the chain, which this new line would have reported as chain BROKEN on perfectly healthy traffic. The owner-only opener above sits inside that section: mode-at-creation and chain ordering are two properties of one write.

Below the shared reporting floor every line withholds its number and says how far along it is instead. A verdict computed over evidence too thin to support it is worse than no verdict, because it teaches the reader to ignore a line that was supposed to be able to say no.

A verdict printed on every exit has to cost what one exit is worth. Each of these lines reads an append-only artifact that grows one row per request and never shrinks — the maintainer's receipts.jsonl is 83 MB — so a render that re-parses them end to end is O(lifetime): a multi-second stall and a memory spike that arrive gradually enough that nobody attributes them to this feature. Three things now bound it. Chain verification streams instead of materialising the chain, and resumes: a small receipts-verified.json records (count, head_hash) and the byte offsets of the last verified receipt, so only rows appended since are re-hashed, and the statement names what it skipped. That checkpoint can only make the answer cheaper, never wronger — it is re-hashed before it is trusted, any failure from a resumed pass is discarded and re-run in full, and a third party handed the file always gets the full pass, as does distil receipts itself — the default is the full pass, and --fast is the opt-in that resumes from the checkpoint. shadow.jsonl is read once per render and shared by every line that quotes it, rather than once per line. And the drift monitor folds only the rows past consumed, carrying the stream fingerprint forward as an accumulator instead of re-deriving it over the whole prefix. On a 200,000-receipt chain and a 50,000-row shadow ledger the exit summary goes from about 1.5 s to 237 ms, with peak allocation a fraction of either file.

The drift state is also written atomically now — temp file, owner-only at creation, then a rename under the same lock. tripped is sticky and capital accumulates across sessions, so a write torn by a crash, a full disk or a Ctrl-C would not have corrupted the alarm noisily; it would have silently reset it to a fresh monitor that reads intact over evidence that said BREACHED. And each verdict is now computed in isolation, which the docstring already claimed: one unreadable artifact drops its own line and the other three still print.

Saying which parts are not on the request path

distil doctor now states plainly that guideline.py's outcome statistics are a zero-sample no-op — record_trajectory_outcome has no callers, so the routing it feeds is running on defaults. Seven research modules (gist, speculative, ensemble, retrieval, output.digest_output_blocks, telemetry.sign/submit, and trajectory_risk.drift_monitor) now open their docstrings with RESEARCH-ONLY — not on the request path. Nothing was deleted and nothing changed behaviour; an inert module that reads as shipped is a claim, and it is now labelled as what it is.

Every receipt now costs one write and one read of the last line, not the chain

Appending a receipt reads the current head hash first — it has to, the chain links each receipt to the one before it — and that read held the same lock the append itself takes, so it serialized every request behind it. It was also an O(chain) scan: head_hash() walked the file from byte zero looking for its last line. On the maintainer's own 94 MB, 102,767-row chain that is 0.53 s spent per request holding a lock every other in-flight request is waiting on, growing without bound as the file does. head_hash() now reads backward from the end of the file in blocks instead of forward from the start, stopping at the first line that parses — a torn trailing line (a write cut short by a crash or a full disk) is skipped exactly as read() already skips it going the other direction, and a line longer than one block just costs another block, never a wrong answer. Chain format, lock scope, and every caller's semantics are unchanged. On a synthetic 100,000-row chain the lookup goes from 325 ms to 0.18 ms.

The receipt chain is sealed into segments, each with a Merkle root an auditor can check alone

The receipt chain is the artifact a team hands to someone who does not trust it, and it only grows: one file, re-hashed end to end to answer anything about any part of it. It is now split. When the active receipts.jsonl reaches the segment size, the append that finds it there (inside the same lock, one stat per request) seals it: a checkpoint is written first — segment id, row count, first and last receipt hash, and a Merkle root over the segment's receipt hashes with RFC 6962 leaf/node domain separation — and then the file is renamed into receipts-segments/. No receipt byte is rewritten, which is also the migration: a chain from before this verifies unchanged and becomes segment 0 on its first rotation. The next receipt links to the sealed segment's last hash, so every segment boundary is a chain link, and a deleted or reordered segment breaks the chain the same way a deleted receipt does.

distil receipts still verifies the whole history, now including every segment against its checkpoint. --segment N verifies one sealed segment against its checkpoint and opens no other file. --checkpoints prints the checkpoint records to pin somewhere outside ~/.distil (each record's sha256 goes to stderr) — whoever can write there can re-seal a segment and its checkpoint together. --prove <request-id> emits an inclusion proof for one sealed receipt; --check-proof re-hashes the receipt from its content, then walks the audit path (RFC 9162 §2.1.3.2), reading nothing but the proof. What a pass proves is exactly what was pinned. --checkpoint-hash <sha256> pins the whole checkpoint record — segment, row count, first/last hash, root — so the output names the receipt's segment and position. --root <hex> alone pins only the tree: it proves membership, and the output names no position, because the row count then comes from the proof unauthenticated and audit paths for different (index, size) pairs coincide (index 1 of 2 also verifies as index 2 of 3). With nothing pinned the proof is checked against its own checkpoint, which is circular; the output says SELF-CONSISTENT ONLY and the CLI warns on stderr. The path length is also checked against the claimed tree size, and every receipt field in a proof is type-checked, so a hostile bundle ("handles": 5) is a clean "malformed proof", not a traceback.

The same type check now runs on every line of the chain, which forced a decision about lines that are not receipts. They are no longer silent. A JSON object that is not a valid receipt ("handles": 5, a string where a count belongs) is BROKEN at that line: no torn write produces a well-formed object, and distil's writer always writes every field with its type, so it can only be an edit — and it stays BROKEN after later receipts chain past it. A line that is not a JSON object at all (a torn trailing write after a crash, foreign text) keeps the existing contract, that it does not invalidate the real receipts around it, but the verdict now reads VERIFIED WITH GAPS — … N lines are not areceipt and were skipped instead of a clean VERIFIED, and the wrap exit line says so too. The resumed (--fast) pass carries that count in its resume point, so it reports the same number as the full pass. A checkpoint's schema field v is now validated as the integer 1; anything else is an unreadable checkpoint, and a checkpoint of another schema is a segment mismatch.

Crash ordering is the design: checkpoint first, rename second. A crash between them leaves a checkpoint with no segment, which every reader ignores and the next seal overwrites, and the active file untouched; a crash after the rename leaves no active file, and head_hash() then reads the newest segment's tail — still a tail read, not a history scan — so the next receipt links correctly. A seal that fails is logged at debug and the receipt is appended to the active file anyway, and this process does not try again for a minute: a seal that keeps failing (an unwritable segments directory; on Windows, a reader holding the active file open) would otherwise re-parse the whole active file on every append, inside the lock every request waits on. The cheap steps that can fail — creating the directory, tightening it to 0700 even if it already existed — run before the parse. verify() runs without the append lock, so a seal landing mid-pass could make a healthy chain read as broken at the segment boundary; a failure is re-run once against a fresh listing when the listing changed, and a real break survives the retry. The resume point behind distil receipts --fast and the wrap exit line records which file its last receipt was in; a rename keeps that receipt's offsets, so the fast pass resumes straight across a seal, and a resume point written before this reads as file 0, which is exactly where the migrated chain lands. Segments and checkpoints are created 0600 in a 0700 directory, like the chain.

One side fix, in the function this rewrote: verify(path) on an explicit file used to write this machine's resume point with offsets from that other file. It no longer does.

Fixed — a claim pointed at the one file that never had its numbers

docs/claims.json's v-cache-aware-vs-naive entry cited docs/CACHE.md for its 33% / 11% / 2× figures. That file contains none of them — the coverage gate's .md skip is why nothing caught it. The figures are real (they are the distil savings --trajectorycorpus/sample_trajectory.json --pricing claude-opus-4-8 table already printed on architecture.html and concepts.html, and the "2×" figure is the separate live-adapter incident documented in docs/cache-contract.html), but no artifact had ever committed the underlying run. benchmarks/results/cache-aware-vs-naive-2026-09-24.json now holds that run's output — produced by executing distil.compress.cache_aware.simulate against the real corpus fixture, not invented — and the entry points at it. tests/test_claims_coverage.py no longer exempts .md artifacts from value-checking; the one other .md-backed entry (the re-read delta's 51.4%, ADR 0010) was checked against this tightened gate and passes.

Fixed — distil discover --since dropped a session whose only recent traffic failed

list_sessions() derived last_ts from booked ledger rows and the manifest's started_ts alone. The ledger only ever gets a row once a request is billed, and started_ts is a session's birth, not its most recent activity — so a session whose only in-window traffic was a failed, unbooked request (bad key, upstream 5xx, client abort) read as stale under --since and was silently dropped from both distil dissect's session picker and distil discover's window, even though it had just been used, and this held even for a session with an old booked row on record: a request from 30 days ago does not make a failure five minutes ago any less recent. Every proxied request appends to sessions/<sid>.requests.jsonl regardless of outcome, so its mtime is now folded into last_ts unconditionally, alongside the ledger and the manifest — whichever source saw the most recent activity wins, rather than the requests file only being trusted when the ledger had nothing at all for that sid. New tests at both the dissect.list_sessions() and discover.scan(since_days=...) layers cover the previously-dropped cases, including the old-booked-row-plus-recent-failure shape.

Fixed — the "2×" figure was carried by a run that never measured it

v-cache-aware-vs-naive above still bundled a "2×" value that its artifact, benchmarks/results/cache-aware-vs-naive-2026-09-24.json, does not state — that figure belongs to a different incident (the pre-1.45 live-adapter cache-bust bug) than the trajectory simulation the artifact actually ran. The artifact itself had also grown a live_incident_2026 block describing that separate incident, which is the same failure mode this fix exists to close: a results artifact must hold only what its run measured. Removed the block, trimmed v-cache-aware-vs-naive's values to the 33%/11% the artifact does state, and moved "2×" to its own entry, u-live-cache-naive-2x, marked unsourced with a check_reason: no committed artifact holds the live A/B's raw numbers, only the narrative writeup on docs/cache-contract.html and its retelling in CHANGELOG.md and benchmark.html. EXPECTED_ENTRY_COUNT in tests/test_site_claims.py moves 51 → 52 for the split.

Fixed — an old booked row could still mask a session's most recent failure

The previous fix folded a session's sessions/<sid>.requests.jsonl mtime into last_ts only when the ledger had no booked row at all for that sid — which missed the more realistic shape of the same bug: a session booked once, long ago, whose only recent activity was a failed request. The old row won the max(), the session still read as stale under --since, and the fix didn't fix the case it was written for. The fold-in is now unconditional — ledger, manifest started_ts, and requests-file mtime all feed the same max(), so whichever source actually saw the most recent activity wins regardless of which of the three happens to be populated. The fixtures that broke under the unconditional rule were pinning their requests file's mtime to wall-clock "now" beside a synthetic historical ts; os.utime'd to match their own fixture clock instead of narrowing the rule to work around them. A new test pins the old-row-plus-recent-failure case.

The output shaper was off, and nothing was deciding

The verbosity directive that shortens replies has existed since output compression shipped. It defaulted to off, and the only way to turn it on was to know the flag — which means the decision about whether to use it was being made by whoever read the docs most recently, not by any evidence. That is backwards for a lossy transform. Every other lossy thing distil does answers to the live paired shadow verdict; this one answered to a default. --shape-output now takes auto, and auto is the default on a metered session. Shaping is on while three things hold and off the moment one stops: the paired decision-equivalence verdict is at or above the one reporting floor the whole product shares (50 A/B, 30 A/A); its harm bound sits inside the pre-registered ±2pp certification budget, read from the same constant distil certify tests against rather than a second number that could drift from it; and the shadow output-token delta excludes zero on the saving side over at least as many samples. Explicit light/aggressive/off still win. Subscription and OAuth sessions are unchanged and always off — that is a policy boundary, not a preference, and an explicit request there is suppressed and says so. The third condition is an enabling signal, not a forecast, and the docs say so where the rule is described: it measures what input compression already does to reply length on your traffic, because a shaping A/B needs a live model and cannot be run at session start. What shaping itself costs is still distil output-savings, after the fact. The gate cannot vote for itself. Once shaping is on, the shadow replay of the served request carries the directive, so "replies got shorter" on those rows is shaping measuring itself — a gate fed them would stay on by construction. Every shadow row now records the levers active when it was measured ("levers": {"compression": …, "shape":…}), and the evidence that turns shaping ON comes only from shape: off rows. Rows written before the tag are excluded too; their shaping state is unknown, not off. The evidence that turns it OFF includes the shaped rows: they are the only ones that measure the directive's own decision-change effect, so once they clear the reporting floor their harm bound must sit inside the same ±2pp budget, or auto resolves off. Without that half, shaping could be turned on but never off. Both sets are read over the last 7 days only (the same recent window distil stats uses), so an "on" decision cannot rest on traffic that has since changed. The tag is per-lever so expand and re-read delta can join it later. Off is never silent. The reason is printed at startup, written into the session manifest beside the level and what was asked for, and read back by distil dissect. The last published live sample (398 A/B, 399 A/A) predates the lever tag and so counts as zero rows; its harm bound of −4.5pp is outside the 2pp budget in any case. auto resolves to off and prints why, which is the feature working rather than a number being withheld. Reply length itself is #185's output verdict line; this change adds no second one.

A digest that stops pointing at what it could just show

A dropped run of lines no longer than the << +N lines, handle=… >> marker that would replace it is now shown inline. The marker format is unchanged and every marker still names its handle; the only difference is that the digest no longer spends a pointer on something cheaper to show. It is more faithful and never more expensive, and the inlined lines are no longer reported to the query flywheel as dropped. The gain is small and it is stated as such. On the corpus it is zero: none of the 51 markers the corpus produces replaces a run that short (benchmarks/first_sight_digest.py, artifact benchmarks/results/2026-09-24/first_sight_digest.json), and bench, verify, validate, retention, fidelity and suite all read exactly as before. It exists for the short gaps between pinned lines that real tool output has and the corpus does not. Two other ways to shrink the markers were measured and not shipped. Dropping the handle from every marker after the first in a block saved a little more, but the distil_expand description tells the model every marker carries one, a kept line that itself contains handle=… could then sit between a bare marker and its real handle, and correcting the description would have cost every user one cache miss on their tools prefix. Shortening over-long head and tail lines to their two ends moved corpus facts from visible to one distil_expand away (retention visible 37.7% to 33.1%) and added a silent failure; a token saved by a round trip is not a token saved.

Why there are no per-command shell profiles

RTK-style profiles — rewriting git status, pytest, ls/grep output with rules per command — were measured before being built, on real traffic only: 3,585 local Claude Code transcripts with the provider's own billed usage, and the original command and output for every shell result. Shell output is 15.8% of that bill (12.6% once whole-file reads, which stay byte-exact, are set aside), and the generic digest already removes 61.5% of those tokens at first sight, cache-stably. Crediting a profile with RTK's best claimed ratio on every output in the families it targets adds at most 1.4% of the bill, because most of the volume is grep, python scripts and sed, not test runners or git. That is below the bar for a second transform surface, so nothing was built. The decision, the numbers and the three conditions that would reopen it are in ADR 0015; the measurement is one command, benchmarks/results/2026-09-24/shell_headroom.py, and writes only its own JSON.

1.53.0 — half of a re-read is a second copy, and a rewritten history is not a cache miss

The through-line: the other end already has the bytes. Inside the conversation, half the mass of a re-read is lines the model is already looking at — a unit nothing here could see, because the near-duplicate gate asked whether two blocks were similar rather than which lines had already been sent. At the provider, a cached prefix was being missed on every single turn by clients that rewrite their own history without changing a token the model reads. Neither change is a new way to compress; both are accounting for what has already been delivered.

The rest of the release closes gaps that share one shape — a rule written for every surface but enforced on one. The quote-hazard counter reported Codex traffic as carrying no edits at all, because apply_patch is a freeform tool with no old_string to read. Vision duplicate elision was Anthropic-only, so an agent looping OpenAI or Gemini paid full price on every repeat of a screenshot it had already sent. Eight file-backed stores fell back to no lock on Windows rather than to an equivalent one. distil wrap reached only agents with an environment-variable routing contract, so Continue, Factory Droid, Oh My Pi and Crush were wrapped by doing nothing — and the registry meant to make concurrent wraps safe collided with itself on Windows' 15.6 ms clock, restoring a config out from under a live sibling. LlamaIndex joins the duck-typed integration set, after the delegating wrapper it started as turned out to fail LlamaIndex's own isinstance checks — which made this module's documented example raise. A benchmark could overwrite the paper's committed source data in place from a smoke run. And the provider-compaction certifier was still scoring itself with the clipped estimator #165 removed from shadow mode; it is now the paired, signed difference it was designed to be all along — no published number moved, but every interval now excludes zero, which the old statistic was incapable of saying.

This one soaks as an rc. 1.52.0 went straight to GA on the maintainer's call and its entry said what that cost. This release does not repeat it. Two of the changes below sit on the request path — the re-read delta rewrites tool results, and prefix replay decides which bytes are forwarded at all — and between them they carry zero hours of real traffic. Every offline gate is green, which is precisely the condition the soak policy exists to distrust: the 1.10.0→1.11.3 day was six releases that each passed review and each broke under use. The soak is the maintainer's own machine under distil wrap, through 2026-09-14. It is also the only instrument that can finish the measurement — the 0%→100% replay table below reports what distil forwards, not what a provider billed. That is the cache-read share in distil dissect, and only a live run closes it.

The provider-compaction certifier was scoring itself with the estimator we deleted

#165 removed a clipped ratio from the live shadow estimator: max(0, 1 - p_AB / p_AA) divides one arm by the other, clamps the result at zero, and therefore prints a perfect score whenever the A/B draw beats the A/A draw by chance. An estimator that cannot come back negative cannot report that compression was harmless — only that it was harmful — so it can never fail.

The same statistic was still running in distil certify-provider, which is the harness behind the Recency Is Not Relevance paper's numbers on Anthropic context editing and OpenAI server-side compaction. It had the defect twice over: the clamp, and a division of an A/B rate computed over fired cases by an A/A rate computed over all of them.

The design was paired the whole time and nobody used it. Every case is observed three times — baseline, a second independent baseline, and the manipulated arm — on the same transcript. So the reported statistic is now the per-case difference

Δ = mean over fired cases of   1{edited ≠ baseline} − 1{baseline′ ≠ baseline}

with a percentile bootstrap 95% interval, unclipped and signed: how much more often the provider's manipulation moved the decision than the model moved it by resampling. The raw A/B and A/A rates are unchanged and now carry Wilson intervals; certification still rests on the distribution-free bound on the raw rate. The old field survives as legacy_adjusted_change_rate so pre-existing artifacts still parse.

No published number moved. The committed run artifacts were re-derived offline from the per-case arm signatures they already record (provider_compaction_to_latex.py--recompute — no API calls, the experiment was not re-bought), and every headline is byte-identical: Anthropic's aggressive configuration still 92.5%, its default policy still 95.0% and 100%, OpenAI still 12.5% and 20.0%, nothing certified at α=0.1. What changed is the column beside them:

Runold clipped ratiopaired Δ (95% CI)
Anthropic, default keep=3, run 2100.0%+97.5pp [+92.5, +100.0]
Anthropic, default keep=3, run 194.9%+92.5pp [+85.0, +100.0]
Anthropic, aggressive keep=091.9%+85.0pp [+70.0, +95.0]
OpenAI, threshold 3k17.9%+17.5pp [+7.5, +30.0]
OpenAI, threshold 1k12.5%+12.5pp [+2.5, +22.5]

Every interval excludes zero, which is a claim the old statistic was not capable of making — and the largest correction, 6.9pp, was in the direction that overstated harm.

A smoke run could overwrite the paper's source data

benchmarks/leave_one_domain_out.py defaulted --out to docs/paper/results/leave_one_domain_out.json, a tracked artifact. A reduced run (--control-reps 5) therefore replaced the paper's E3 numbers with a smoke test's, in place, silently. It was caught by reading a diff, which is not a control.

Every benchmark that can reach a committed results file now defaults to the git-ignored benchmarks/results/scratch/ and prints where it wrote. Publishing takes --write-tracked or spelling the tracked path out in full. That covers leave_one_domain_out.py, trajectory_certificate.py, trajectory_bound.py, skeleton_certificate.py, and the E7 runners swe_bench_e2e/{sample,aggregate,preload_images,run_agent}.py. A test greps benchmarks/ for writers that reach docs/paper/results/ and fails on any that is not guarded, so the next script to reach for a tracked default trips CI rather than a reviewer.

The second read of a file need not re-send the first one's lines

The exact-quote guarantee keeps file content byte-exact forever, and that is expensive on exactly the traffic it protects: under a client that caches its whole history — which is what Claude Code does — a superseded read cannot be demoted either, so the bytes stay. The same investigation measured what those bytes are.

Read/cat results that re-read a path already read this session51.4%
of those, byte-identical to the earlier read12.8%
of a re-read's tokens, lines already delivered verbatim earlier in the conversation50.6%

Half the mass of a re-read is a second copy of lines the model is already looking at, and neither existing mechanism could see it. Exact dedup needs identical bytes, and 87% of re-reads are not identical. cachedelta's near-duplicate gate compares whole blocks at a 0.5 difflib ratio and misses 608 of 726 changed re-reads, because two disjoint windows onto one file are not 50% similar as blocks even when every line they share is byte-identical. It fires on 5.2% of re-reads and removes 0.23% of tool-result mass.

The unit was wrong. Block similarity asks "is this nearly the same document"; the useful question is "which of these lines have I already sent". So the new transform matches on line content: the longest run a read shares with an earlier read of the same path becomes a reference stub, every other line stays verbatim, and the elided lines come back byte-exact from distil_expand like every other distil stub.

«distil-reread handle=a1b2c3d4» lines 1-20 of this result are byte-identical to
lines 41-60 of the earlier read of /app/handlers.py, which is still in this
conversation verbatim. Call distil_expand with this handle to recover them here.

Why this does not weaken the guarantee. The promise is about the conversation, not about any one block: an Edit(old_string=…) applies if its quote occurs byte-exact anywhere in the payload distil forwarded, and an agent quoting lines it saw can be served by either copy. Six rules keep that true.

  • Only a read matched by the tool-NAME table may be referenced. Read, view, read_file and friends are exempt unconditionally, so the referenced lines cannot leave the conversation later. A shell cat may be a target but never a base: its exemption is conditional on not being superseded, and a whole-file shell re-read supersedes exactly the block it would want to point at. Pinning the base to stop that would forfeit the base's whole digest to save the same bytes on the copy — a wash at best.
  • References never chain. An elided block is not itself a base, so every stub points at literal bytes.
  • Cuts interior to the block are pulled back 20 lines. A quote inside the elided run survives in the base and one outside it survives here; only a quote straddling a cut is in neither contiguously, so the cut is moved far enough that it would have to overhang by more than twenty lines to break. A run reaching the block's own first or last line takes no margin there — the agent saw nothing beyond it in this block.
  • Runs under 8 lines are left alone, so a coincidental match on blank lines or a repeated return None never produces a stub.
  • Lines carry their terminators into the comparison. Stripping them would call a CRLF read and an LF read of one file identical, and eliding the only LF copy while the base holds CRLF bytes makes the stub's claim of byte-identity false.
  • The freshest tool output is never elided. The recency carve-out applies here like everywhere else: a re-read the agent has just issued is exactly the output it reasons over to choose its next action. The plan is computed from the prefix regardless — which blocks may serve as bases never depends on the sliding window — and only its application is gated, so nothing the provider has cached moves.

Why it is cache-safe without a volatile-suffix gate. The plan is a pure function of the message prefix — what a block encodes to depends only on the blocks before it — so a block's bytes never change across turns at all. That is stronger than the contract asks for, and it is the only construction that works here: ADR 0008 records that under the Claude Code shape the whole history is cached and there is no uncached tail, so a transform gated on the boundary would never run. Worse, such a gate flips a block from stub to verbatim exactly as the boundary advances past it — the failure the contract exists to catch. The cache contract gains clause (e), tests/test_reread_delta.py asserts byte-stability under the moving-marker shape, and ADR 0010 states the trade.

Measured. On the codebench corpus (read → edit → re-read, 20 sessions / 320 turns) under the client shape that bills — newest turn pinned, whole history cached — the PAYG digest row moves from 31.4% to 41.7% token savings and 35.8% to 49.1% cache-aware dollar savings. Before/after output in benchmarks/results/2026-09-06/. distil bench is unchanged byte for byte: its corpus has no re-read of one path through a name-keyed read tool, so the delta never fires there.

And codebench is nothing but reads, so it overstates. On real traffic name-keyed reads are 10.7% of tool-result mass, about half are re-reads, and about half of a re-read's tokens were already delivered — a live ceiling near 2.8% of tool-result mass. That is the number to quote.

One benchmark caveat this surfaced, stated rather than hidden. benchmarks/codebench.py builds its sessions with no cache_control marker and then prices them with a cache — the longest stable prefix is billed at the cache-read rate. No Anthropic client looks like that: Anthropic caches only what the client marks, so an unmarked request has no cached prefix at all, and every recency-anchored carve-out distil has is charged there for busting a cache that was never created. Under that shape the digest row reads 7.1% tokens for −15.4% dollars. The artefact predates this release — distil-verbatim shows 18.5% tokens for 3.7% dollars on the same corpus — and was invisible only while the digest row was 0.0% and nothing moved. New runner benchmarks/codebench_marked.py replays both shapes so it is reproducible rather than asserted. The corpus itself is left alone: changing it would move the competitor rows on docs/compare.html too.

distil validate gains four re-read shapes under the existing quote-survival invariant — a re-read at a different offset, a quote straddling a cut, read → edit → re-read, and a whole-file read after a partial one — plus a test that fails if none of them still produces a stub, so a silently-dead transform cannot pass as a silently-safe one. On an observed quote miss the widen reaction now turns the re-read delta off for the rest of the session alongside dropping supersession. The planner costs 0.053 ms/turn.

The quote-hazard counter now covers Codex

1.51.0 shipped the exact-quote exemption on all three adapters but the measurement on the Messages path only, on the reasoning that Edit/MultiEdit live there, and listed the Responses API as a known follow-up. Codex does the same edits under different names, so Codex traffic was reported to distil dissect as carrying no edits at all.

apply_patch is the gap. On GPT-5 models it is a freeform custom tool: per its own Lark grammar the argument is the patch body itself, not JSON, so there is no old_string field to read. What has to match the file is each hunk's pre-image — the context and removed lines in order — and it runs through the + lines, since an added line is not part of what the patcher is looking for. provenance.patch_quotes extracts exactly that; response_edit_quotes normalises both the custom_tool_call and function_call shapes. A one-line hunk counts only when the line is long enough to discriminate: single-line hunks are common, so dropping them all would blind the gauge exactly where it is meant to see Codex, but a one-line quote of } occurs in any payload and would be booked as a survivor on a request where nothing survived. Under-counting is silence; over-counting is a false all-clear on a safety gauge. observed_view now excludes model-output item types (function_call, custom_tool_call, reasoning, …) as well as assistant messages — without that the patch envelope quotes the file to itself and the check could only ever pass.

The Responses path gets the same widen-on-miss reaction as the Messages path. Nothing changed downstream: both adapters already share one thread-local counter, so distildissect reports Codex through the line it already had.

The cached prefix now survives the client rewriting its own history

ADR 0008 promised byte-stable forwarding when the client re-sends a message byte-identical, and clause (d) said a client that rewrites its own history gets no promise. That clause was written as an edge case. It is the common case: agentic clients rewrite their history on every turn without changing a token the model reads — the cache_control breakpoint advances to the newest block, an SDK shim stamps positional index fields, a string becomes a single text block and back again. distil forwards what it receives, so all of it reached the provider as changed bytes and missed a prefix the provider was already holding, at the write rate.

Measured offline against distil's own adapters (benchmarks/prefix_replay_stability.py, 8 turns; share of re-sent messages forwarded byte-identical to the previous turn):

provider shaperewritebeforeafter
anthropic/messagesindex stamped0.0%100.0%
anthropic/messagesstring/block sugar0.0%100.0%
openai/chat-completionsindex stamped14.6%100.0%
openai/chat-completionsstring/block sugar0.0%100.0%
openai/responsesindex stamped0.0%100.0%
openai/responsesstring/block sugar0.0%100.0%
gemini/generateContentindex stamped0.0%100.0%

Zero, not merely degraded. The un-rewritten control reads 100% before and after, so replay repairs rewrites and does nothing when there is nothing to repair. Added latency is ~0.13 ms per request. This is what distil forwards, not what the provider billed — that is the cache-read share in distil dissect, and it needs a live soak.

  • Forwarded-bytes prefix replay (distil/prefixreplay.py, ADR 0011). Per conversation lineage — leading system-run + model + tools + the canonicalised head, per session — distil remembers the previous turn's (original, forwarded) pair. For the longest canonically-equal leading prefix it forwards the bytes it forwarded last turn, re-placing the client's current cache_control markers at the client's current block positions, and compresses only the divergent suffix. Where a marker has nowhere to sit on last turn's form — the client re-spells a bare string as one text block and puts its breakpoint on it — replay stops there rather than forward a form with the breakpoint missing: on Anthropic the breakpoint is the cache entry, and dropping it would buy the hit by destroying the thing being hit. On by default; --no-prefix-replay opts out on both wrap and proxy. State is in-memory, LRU-bounded, and deliberately not persisted: it holds message bytes, and the TTL'd restore store is the only place content is allowed to rest on disk. Fail-open — any exception forwards exactly what the compressor produced.
  • The comparison ignores four things and nothing else: cache_control, index, the interchangeable spellings of a single text block (bare string, text, input_text, output_text), and JSON key order. Tool payloads are opaque — the arguments going out and the result coming back, compared as they arrive and never structurally rewritten, because those rules are about message structure and none of them hold one level down: a payload can carry a key called content (folding it would call two different tool calls equal) or a key called index that is data. Gemini's functionResponse.response is arbitrary tool output, and stripping index inside it would forward one result's bytes for a different result. citations/annotations are deliberately left out: the cost of excluding them is a missed hit, the cost of including them wrongly is a wrong prefix.
  • Replay restores bytes, never decisions — the clause that makes it safe, and not the obvious design. distil's compressor is not a pure function of one message: the exact-quote guarantee keeps a tool result verbatim because of an Edit that arrives later in the list, so a block legitimately digested at turn N can need to be verbatim at turn N+1. Overlaying turn N's stub there would break the agent's next edit to buy a cache hit. So a message is replayed only when this turn's forwarded form is itself canonically equal to last turn's, which makes "replay never changes semantic content" the loop condition rather than a claim. distil validate asserts it as a sixth invariant anyway, because a loop condition is what a later optimisation deletes.
  • Proof. tests/test_cache_contract.py gains the three rewrite shapes × four request shapes, each with a control arm: the test fails unless the rewrite demonstrably shortened the byte-stable prefix without replay and demonstrably did not with it — a stability assertion whose control never destabilises proves nothing. Plus: a semantic edit breaks the prefix at exactly its index; a changed tool_result is never overlaid with older bytes; a re-signed thinking block breaks the prefix while an unchanged one is replayed byte-for-byte; the lineage never crosses a model/tools/system change; state is bounded and an oversized history is released rather than held; and the two transforms that share the cache contract are driven together through the real proxy — a re-read delta stub (ADR 0010) is forwarded byte-identical on the next turn, and when a later Edit makes the compressor withdraw that stub the verbatim block goes out instead of last turn's bytes. Replay never writes back into the state it replayed from, so a concurrent request on the same conversation cannot be handed this one's marker placement. The canonicalizer's strip set is pinned from both sides — a dropped else in its recursion once made every message canonicalise to {}, and the parametrized tests caught it only by luck.
  • Measurable. The per-request record gains content-free replay_hits / replay_misses / replay_restored, and distil dissect reports them per session beside the cache-read share. restored is never folded into hits: hits with zero restored is the healthy steady state, and adding them together would make a dead mechanism look identical to a working one. A session run with --no-prefix-replay records no counters rather than zeros, and dissect prints not recorded: zeros there would report a switched-off feature as one that ran and held nothing, directly beside the cache-read share you would be using them to explain.
  • All three servers get it, through one shared prefixreplay.apply at the same point in each: the threaded proxy, the async proxy (--async, which now honours --no-prefix-replay too), and the multi-tenant gateway. A default-on feature that only one server has is a feature its users do not have — the 1.46.0 lesson, since managed installs run distil proxy. A server that re-serialises every body still benefits: its output is deterministic given the same items, so replaying the items is what makes its prefix stable. Each scopes the lineage by whose credential it is, because a cached prefix belongs to one credential and the lineage key is otherwise content-derived: the gateway per tenant, and the plain proxies by a hash of the client's own API key, which is the only identity they have. Asserted on both by a test that posts the identical conversation under two identities and fails if the second is served the first's bytes. All three carry --no-prefix-replay, including the gateway, whose chart exposes it as gateway.prefixReplay — the chart is where a default-on feature silently becomes default-off, so the value and the arg that reads it are both pinned by a test.
  • A block whose cache_control marker did not move is returned untouched rather than rebuilt. Rebuilding moves the marker to the end of the key order, and JSON key order is part of the bytes the provider hashes — so the naive marker re-placement busted the prefix the first time it replayed a block whose marker was not already last. The per-server tests caught it; the adapter-level ones could not.

Windows ran every file-backed store with no lock, and two adapters saw every repeated image at full price

Two platform gaps, closed together because both are "works on Linux, quietly worse on Windows" — the kind of defect a Linux CI matrix cannot see.

  • File locking degraded silently on Windows. gateway_keys.py, audit.py, ledger.py, census.py, shadow.py, retention.py, surfaces.py, and mcp_server.py each guarded import fcntl with a bare except ImportError and fell back to no lock at all rather than an equivalent one — a concurrent-writer race that Linux/macOS CI never exercises because fcntl always imports there. _filelock.py is now the one place any of them ask for mutual exclusion: fcntl.flock on POSIX, msvcrt.locking against a <path>.lock sidecar on Windows, fail-open (no lock, not a crash) if the platform call itself errors. All eight call sites now route through it, and the tests that were skipped under win32 run unconditionally by exercising the Windows branch directly rather than requiring a Windows runner.
  • Vision duplicate elision (ADR 0003) was Anthropic-only. The certificate-gated dedup that replaces a repeated image with a reversible reference — same first-occurrence rule, same "a URL is never proof of identical bytes" rule — only ever saw Anthropic's image blocks. OpenAI's image_url (Chat Completions) and input_image (Responses API) parts, and Gemini's inlineData/fileData parts, passed through untouched, so an agent looping OpenAI or Gemini paid full price on every repeat of a screenshot it had already sent. The identity/elision logic now lives once in compress/vision.py (elide_or_keep), shared by all three adapters behind the same certificate gate and the same image_kept/image_elided census reasons distil dissect already reports. proxy._count_messages gained the matching image_url token count so the before/after savings are on the same scale as the new elisions — without it, an elided OpenAI image reported zero savings despite being genuinely removed.

wrap reaches config-file-only agents, and a LlamaIndex integration

distil wrap previously covered only tools with a documented environment-variable routing contract (AGENT_PRESETS). Continue, Factory Droid, Oh My Pi, and Crush have no such contract — the base URL lives only in a config file — so wrapping them silently did nothing. A second registry, distil.config_wrap.CONFIG_PRESETS, adds four strategies for this shape: cn (Continue) generates a session-only temp config passed via --config, so nothing stable is ever touched; droid (Factory Droid) JSON-merges a distil entry into the local-override layer settings.local.json, byte-exact backed up and restored if the file already existed, created and deleted if it didn't; omp (Oh My Pi) splices a marker-fenced block into models.yml with the same backup/restore contract; crush (Crush) JSON-merges a distil provider entry into the legacy crush.json — Crush's current config format is a Bash script (crushrc), whose own docs call the JSON format deprecated but still read, one tier below crushrc. All four restore on normal exit, on SIGTERM/SIGHUP (the existing signal-to-KeyboardInterrupt path), and — for the case nothing can catch, SIGKILL or power loss — restore_stale_backups() sweeps every registered path at the top of the next distil wrap and repairs it before that session starts. No credential is ever invented: a preset that needs an API key skips silently (touching nothing) when that key isn't set, matching AGENT_PRESETS' extra_env rule.

The one crash a backup-and-restore contract cannot survive is a crash before the backup exists, and the session registry that lets concurrent wraps share one config opened exactly that window: "do I write the backup?" was answered by "was the registry empty?", and the registry entry went in first. A SIGKILL between the two left a registry holding one dead entry, so the next distil wrap read "not first", skipped the backup, and patched the real config anyway — an injected distil block with nothing on disk able to undo it, duplicated on every subsequent run for the two splice-strategy tools. The claim, the backup and the patch are now one critical section, in that order, under one distil._filelock lock (the cross-platform helper from the entry above, replacing this module's own fcntl/msvcrt shim), and the backup is created based on whether the backup file exists — so a crash at any point either leaves the config untouched or leaves the recovery material behind. Every patch is idempotent besides: re-patching a config that already carries a distil entry replaces it rather than appending a second one. restore_stale_backups() also sweeps a registry directory whose registrants are all dead but that has no backup beside it — a dead session's bookkeeping, which used to claim the path forever.

Two more from the same review, both Windows-shaped. _pid_is_alive called OpenProcess through ctypes without pinning a restype, so ctypes' default c_int truncated the 64-bit HANDLE on Win64: GetExitCodeProcess/CloseHandle then failed with ERROR_INVALID_HANDLE, the fail-safe fired, and every pid read as alive — crash recovery silently inert on the platform this registry exists for. The three signatures are now pinned explicitly and asserted by a test, because the fake-kernel32 test cannot catch an ABI bug: a Python stand-in has no ABI. And the liveness check caught only OSError, while a platform that is neither POSIX nor real Windows raises AttributeError from ctypes.windll — an exception at the top of distil wrap that wouldn't degrade crash recovery but stop the CLI from starting. Both it and restore_stale_backups() are now fail-open per path: a registry in any shape at all costs you crash recovery for that run, never the wrap.

A third one sat underneath those, and it is the one that made concurrent config-file wrap genuinely wrong on Windows. A session's registry entry was named <pid>.<time.monotonic_ns()>, which assumes the clock advances between two claims. It does not on Windows, where 3.12's monotonic clock ticks about every 15.6ms: two sessions starting inside one tick got the SAME filename, the second's touch() landed on the first's entry, and whichever exited first unlinked the only entry, concluded it was the last session, and restored the config out from under its still-running sibling. The entry is now named and created in one step by tempfile.mkstemp, so the filesystem guarantees uniqueness rather than a clock. The regression test installs a clock that never moves at all, which reproduces the Windows-only failure on any platform.

Amp, Mistral Vibe, and OpenClaw were investigated and are deliberately not included. Mistral Vibe's config shape could not be verified against an authoritative source. Amp's amp.url setting is real and current, but it belongs to the VS Code extension's own settings schema, not the standalone CLI wrap would launch — that CLI's settings reference has no base-URL key at all, and documents HTTP_PROXY as its own traffic-redirect mechanism, a different shape than the rest of this registry. OpenClaw's models.providers.<id>.baseUrl is verified but belongs to a persistent multi-channel gateway daemon, not the one-shot session distil wrap models — wrapping it would misrepresent what the flag does.

distil.integrations.llamaindex adds LlamaIndex to the duck-typed, dependency-free integration set alongside AutoGen and Agno: DistilNodePostprocessor compresses retrieved node text (drops into node_postprocessors=[...], no subclassing required — LlamaIndex calls postprocess_nodes/apostprocess_nodes directly, with no isinstance check), DistilLLM wraps an LLM so outgoing chat/complete calls (and their async/streaming siblings) are compressed transparently, and compressing_tool wraps a plain callable for FunctionTool.from_defaults(fn=...).

DistilLLM hands back the LLM re-typed as a transparent subclass of that LLM's own class rather than a wrapper object around it. A node postprocessor really is only duck-typed, but an llm= argument is not: measured against llama-index-core 0.14.24, resolve_llm() — behind both index.as_query_engine(llm=...) and Settings.llm = ... — ends in assert isinstance(llm, LLM), and every Pydantic component with an llm: LLM field (FunctionAgent among them) rejects a non-instance outright. The delegating wrapper therefore failed this module's own documented example. Registering as a virtual subclass is not a fix either — Pydantic disables register()-based isinstance support and warns that it does. The module still imports nothing from llama_index, and anything whose class can't be subclassed falls back to plain delegation.

1.52.0 — the guarantee covered the wrong half, and the estimator could not say no

The through-line: distil's guarantees were narrower than the traffic they claimed to cover, and the instruments pointed at them could not report bad news. The exact-quote rule keyed on tool names and missed three quarters of the file reads on real traffic; the Responses API was compressed by one server of three; the shadow estimator was clamped so it could print 100% and could never report harm. Each of those is now the size it says it is. The proof surface grows a written cache contract, an adversarial suite, and a measured degradation curve.

This release skips the rc soak. RELEASING.md requires any release that changes runtime behaviour to bake as an rc for at least three days of real work first, and this one changes the request path substantially: all three server adapters, the compressor's provenance rule, the shadow sampler. The maintainer's decision is to ship on the gates instead — the suite, the coverage floor, bench, verify, validate --adversarial, the retention and fidelity harnesses, and the certificate. What that costs is worth stating rather than hiding: these paths carry no hours of real use, and the failure mode the soak policy exists to catch is precisely the one that looks correct in review. The 1.10.0→1.11.3 day is the precedent.

The exact-quote guarantee now covers the shell, where most reads happen

1.49.0 stopped distil from digesting file content the agent has to quote back byte-exact in an Edit(old_string=...). It keyed that exemption on the tool name: Read, Grep, Glob, view, str_replace_editor and the common MCP filesystem equivalents. Correct rule, wrong half of the traffic.

Measured over 2,489 local Claude Code sessions (52,937 tool results, 19.4M tool-result tokens, content-free):

share of tool-result mass
reads matched by the 1.49.0 name table10.7%
whole-file reads issued through the shell (cat, head, tail, nl, sed -n '1,40p')33.6%

Three quarters of the file reads on real traffic go through the generic Bash tool, and every one of them was still being digested. Of 3,445 Edit/MultiEdit quotes traced to a tool-result source, 39.3% had no surviving byte-exact copy in what distil actually forwarded — and 629 of those 726 failures had a shell source. That is the same failure 1.49.0 was written to stop: the edit cannot match, so it silently does nothing, and the run ends with the agent reporting success and the disk untouched.

Provenance is now read from the command, not just the tool name. A shell call whose every stage is a whole-file reader, with no pipe and no redirection, produced file bytes and is exempt. The classifier is deliberately narrow, because the question it answers is narrow — are the bytes the agent saw a verbatim slice of a file?

  • cat a b, head -n 50 f, cat -n f, sed -n '1,20p' f — yes.
  • cd /repo && cat main.py — yes, and this one matters most: it is the commonest way an agent reads a file, so a rule requiring every stage to be a reader would refuse it and digest the quote. What the agent quotes from is what the last stage printed.
  • cat f | grep x — no. The agent saw grep's output, not the file's.
  • cat f > out, … | tee out, cat << EOF — no. Those bytes are not what came back.
  • sed 's/a/b/' f — no. That is a transform.

Only the latest read per (path, slice) is kept: a superseded read is one the agent has a fresher byte-exact copy of. Slice, not path — sed -n '1,80p' app.py and sed -n '200,280p' app.py are two disjoint windows onto one file, and treating the second as a fresher copy of the first digests 80 lines the agent may still have to quote. A later read supersedes an earlier one only when it covers it: the identical slice again, or a whole-file read.

A numbered listing is not a byte-exact copy of the file, and the first cut of this classifier read cat -n app.py as a whole-file read like any other — so it superseded an earlier plain cat app.py and digested the one copy the next Edit(old_string=...) had to match. nl, cat -b, cat -A/-v/-E/-T, cat -s, less -N, head -v and bare bat all rewrite the bytes they print. Each is now tagged decorated from a per-command flag table, and a decorated read covers nothing but an identical decorated re-read — in both directions, since plain output is no superset of numbered output either. It stays exempt itself, so nothing the agent saw is lost; it just cannot stand in for the file. Flags that only change buffering (cat -u) are still a whole-file read, and bat --plain prints the file's own bytes so it supersedes normally. The quote guard did repair this after the fact, but only once an Edit was already in the history — the request that digests the read is sent before the model writes that edit.

Supersession also stops at the client's own cache_control breakpoint, because flipping a block the provider has already cached from verbatim to digest rewrites the entire prefix at the 1.25x write rate. That is the trade the recency carve-out was re-anchored to avoid, and it was measured at 2x the cost of compressing nothing at all.

What this costs, stated honestly. The investigation projected 7.7pp of tool-result mass for a latest-per-path rule, against a 27pp upper bound for exempting every whole-file read. That 7.7pp assumes a superseded read can always be demoted. Here it cannot: under a client that marks most of its history cacheable — which is what Claude Code does — the cache gate means supersession rarely fires, and the exemption degrades toward keeping every whole-file read. So the forfeited digest saving in practice sits much closer to the 27pp figure than to 7.7pp.

That is the intended trade, not a regression. There are three ways to handle a read the agent may quote: digest it and break the edit, demote it and rewrite an already-cached prefix at the write rate, or keep the bytes. Only the third costs nothing but tokens, and the tokens it costs are billed at the cache-read rate rather than re-written. Exempting every whole-file read outright would still be worse — it buys 187 more quotes for 3.8M more tokens with no cache benefit — so the slice-aware rule is kept.

The rule lives in one place (distil/compress/provenance.py) and all three adapters use it. Chat Completions, the Responses API and Gemini generateContent had no exact-quote exemption at all before this — Codex and Gemini agents were exposed to the original 1.49.0 bug in full.

--session-delta voided the guarantee outright. delta_encode ran before compression and replaced tool-result bodies with delta references, including the reads the compressor was about to keep byte-exact. A reference is not a copy. The order stands — running delta after compression would delta against digest stubs and rewrite already-cached blocks, breaking the cache-monotonicity that suffix-only encoding exists to preserve — and delta now skips the exempt blocks instead. They stay registered as delta bases, so later re-reads still dedup against them.

Two census buckets, not one: tool_result_exact_quote is what the 1.49.0 name rule froze, tool_result_shell_read is what this one adds. distil dissect sums them by reason with no change on its side, so the new rule's cost in production is a number you can read rather than one you have to infer.

The guarantee now measures itself. Every request records, content-free, whether each literal-match edit's old_string still occurs in the payload distil forwarded: two counts, no text, surfaced in distil dissect and as an anomaly line when any quote is lost. On a miss the whole provenance class stops digesting for the rest of the session. distilvalidate gains a sixth invariant, quote-survival, over synthesized read-then-edit sessions (Read, cat, head -n, sed -n, cat -n, cd … && cat, a re-read, and a read whose content is shaped like a digest stub). Nothing gated this before, which is why 1.49.0 shipped half-covered and nine releases went by without noticing.

Known follow-up: the quote-hazard counter and the widen-on-miss reaction run on the Messages path only, where Edit/MultiEdit live. The other two adapters get the exemption but clear the counter rather than compute it, so a stale count is never reported against them. Wiring it through the Responses API, where Codex's str_replace_editor lives, is not done here.

benchmarks/codebench.py is relabelled, not re-measured. It is an upper bound on what the guarantee costs: ~55% of its tool-result tokens are file reads that must stay byte-exact and the rest is too short to digest, so distil's 0.0% digest row there is the policy working, not the compressor failing. The share is now printed under the table, computed with the live classifier.

Shadow mode: an estimator that can say no

Live shadow, build 1.51.1, on the maintainer's lossless-only traffic: 44 A/B and 11 A/A samples. Raw agreement 81.8% [67.3, 91.8]; the model's self-agreement on identical input 84.8% [71.8, 92.4] over 46 byte-identical replays. Statistically indistinguishable (Fisher p=0.63) — and 32 of the 44 "A/B" samples had no bytes changed by compression at all. Below the reporting floor, and not a verdict. Reading that sample end to end turned up an estimator that could not have told us otherwise.

The arms were not what they said they were. _serialize_if_changed returns the ORIGINAL bytes when no transform fired (deliberately — re-encoding busts the prompt cache), but the sample was still labelled by an independent 1/3 coin. So the two arms were the same bytes and the row measured the model against itself, filed as evidence that compression is safe. Identical bytes are now reclassified at spawn time as an A/A sample, whatever the coin said; it also costs one fewer replay. The A/A arm itself replayed the COMPRESSED body twice — that is a B/B arm, which under digest mode measured self-agreement on the wrong distribution entirely and could not detect a compressor that changed every decision consistently. A/A now replays the original. And the call order was fixed, so the warm-prefix-cache read always landed on the same arm; the replays are now issued in shuffled order.

The estimator has an interval, and is allowed to be negative. The old number was p_BA / p_AA between two disjoint request sets, clamped with max(0, ...). It printed exactly 100% whenever the A/B draw beat the A/A draw by chance, and it could never report harm. The paper defines a paired difference (§ "A/A self-agreement control"); the serving path now implements it. Each sampled request is replayed three times — A and A' on the original context, B on the compressed one — and the row stores 1{A==B} − 1{A==A'}, reported as a mean difference with a bootstrap 95% CI, unclipped, alongside raw p_AB and p_AA with Wilson intervals. That is 1.5x the replay budget on samples where compression actually changed something, and only there; DISTIL_SHADOW_PAIRED=0 restores the cheaper two-replay design (with the A/A arm fixed). SIG_VERSION goes to 5 — the rule on that constant is to bump on any change to how the compared sample is generated, and v4 rows are not comparable to these. Old rows stay readable under --all, labelled legacy-unpaired.

One reporting floor, everywhere. There were three. VERDICT_MIN_AB/ VERDICT_MIN_AA (50/30) gated the status line and the proof ledger; shadow-stats printed at 25/10; the census feed and the web dashboard published as soon as any A/A baseline existed (n≥10). The lowest of the three fed the adoption page's decision-equivalence ring, so the most public claim rested on the weakest evidence. Every surface now uses the same pair and says below reporting floor (n=…) otherwise — and the adoption caption states the estimator rather than the word "provably". The sampling diagnostics still print below the floor, because a thin sample is usually a replay that is failing.

Two things the reports were quiet about. lossless-only and digest measure different things; pooled, they report the average of two experiments nobody runs, so each row now carries its mode and shadow-stats breaks them out. And force_deterministic only pins temperature 0 when the body already carries the field and thinking is off — on Claude Code traffic neither holds, so the replays sample hot. Rather than fake determinism, each row records whether it was pinned and shadow-stats prints what share of the sample ran hot, and why the A/A arm is what absorbs it.

One consequence worth knowing: on lossless-only traffic, where compression often changes no bytes at all, those samples now count toward the A/A arm only — so the A/B floor takes longer to clear, and shadow-stats says how many samples that was. The surfaces stay blank meanwhile. That is the honest state, not a regression.

Output-token accounting. The paired replays already return both responses, so the provider's own usage is free to record. Each row now carries input and output tokens for the A and B arms, and shadow-stats and distil dissect report the mean output-token delta with a CI plus the net dollars — input saved at input prices minus extra output at output prices. Output is priced several times input, so compression that makes the model answer at greater length can cost more than the prompt it shortened; nothing in distil could see that before. Cache fields are deliberately not priced in: the two arms hit the prefix cache differently by construction.

Provider parity across all three servers

The Responses API was compressed by the sync proxy and forwarded raw by the async proxy and the gateway; OpenAI Chat Completions was injected with Anthropic's expand-tool schema, which that provider rejects; the Anthropic SSE splice ran on OpenAI and Gemini streams and corrupted them; Azure OpenAI's paths matched nothing anywhere. The aider and grok wrap presets set environment variables those tools never read, so both routed nothing while reporting success.

The gateway no longer emits a digest stub it cannot restore. It injects no distil_expand tool and runs no expand loop, so Tier-1 there was irreversibly lossy on every pay-as-you-go session — the marker named a recovery that did not exist. It is Tier-0 only until it grows the expand loop, which lowers its savings on unstructured content and makes what it does report honest.

The cache contract, written down and made falsifiable

Prompt caching is the largest term in the cost model and it fails silently: rewrite one byte at or before the provider's boundary and the entry for the whole prefix is discarded, the request still succeeds, and the only symptom is the bill. distil has been on the wrong side of this once already — a recency window counted back from the end of the message list produced zero cache reads and 2x the cost of compressing nothing.

ADR 0008 states four clauses: same bytes in means same bytes out at or before the boundary; compression touches only the volatile suffix; digests are deterministic because handles are content-addressed; and what is explicitly NOT promised (a client that rewrites its own history, provider TTL, and the last-k window when no marker is sent at all). tests/test_cache_contract.py replays growing six-turn sessions through the same public entry points the proxy calls, for all three provider shapes and three Anthropic marker placements, using a high-water mark rather than the current turn's boundary — once the provider has cached through an index, rewriting it later invalidates the entry. Bodies are compared as the proxy serializes them, not in a normalized form: key order is part of the prefix, so a transform that rebuilt a dict in a different order would bust the cache and still pass a sort_keys comparison.

The contract holds; no bug was found. That is worth a test, because the property was true and undefended. dissect grows a cache-read share line so the contract can be checked against what the provider actually did — and it renders None, never 0.0, when the fields are absent. "Not measured" and "never hit" are opposite diagnoses.

Compression as an attack surface

COMA (arXiv 2510.22963, ASE 2026) shows an attacker who controls untrusted input can perturb it so the compressor discards task-critical content. The agent then acts on a context missing the line that mattered, and nothing reports a fault: the request succeeds, the savings look good, the certificate is unaffected, the answer is wrong. This lands harder on distil than on a summariser, because the keep policy is legible — a heuristic anyone can read is a heuristic anyone can bait.

So ADR 0009 declines to claim the keep policy as a security boundary. Reversibility is the boundary. The invariant asserted against hostile input is a disjunction: the load-bearing line survives in what distil forwards, OR it was folded, the block is reversible through a handle distil issued, and the stub declares the elision. An attacker can push a line from the first branch to the second; they cannot push it out of both.

distil validate --adversarial adds seven cases — decoy verdict flooding, dedup baiting, salience baiting, handle forging, budget starvation, expand-tool injection, and cross-block starvation. 102/102 checks pass across 19 cases, with two honest findings pinned by tests rather than smoothed over:

  • dedup baiting is a real hit. Shape-based dedup normalises digits away, so 300 attacker lines differing from the genuine error only in a shard number collapse with it and the real line IS folded. Reversibility saves it. Its test fails if the behaviour changes in either direction, so an improvement stays deliberate.
  • decoy flooding costs savings, not correctness. It drives a block to exactly 0.0%. Correct trade, real cost, stated as such — denial of savings, not denial of the answer.

The paper's mitigation (trusted/untrusted budget isolation) already holds here by construction rather than by policy: there is no global keep budget anywhere. That is a claim about an absence, and an absence is what a plausible future "keep top N lines per request" optimisation would quietly fill in — so it is asserted as an equality. A trusted block must compress to exactly the same bytes with or without a 4000-line attacker block beside it, in either order.

distil bench --curve — the shape of the tradeoff, not just the verdict

distil bench answers "is the shipped strategy non-inferior?" with a yes. That is the right question for a gate and the wrong one for a decision: it says nothing about the shape of the tradeoff on either side of the operating point. Anyone choosing between lossless-only and the digest, or wondering what the aggressive rung would buy, has had no measured answer.

distil bench --curve measures every rung over the offline corpus — token savings, fact recall from the retention harness, reversibility, and latency. No API calls, no network, no model, so it runs per-commit rather than once a quarter. Results go to benchmarks/results/curve.json and a stdlib-generated chart to docs/assets/curve.svg, both stamped with the version and date they came from.

Measured on the 9-trajectory corpus at 1.51.1:

rung             savings   recall  visible  lost  reversible
none                0.0%    1.000    1.000     0     yes
tier-0 only         0.0%    1.000    1.000     0     yes
subscription        0.0%    1.000    1.000     0     yes
lossless           47.0%    1.000    0.769     0     yes
aggressive         65.5%    0.551    0.551   828     NO

Three things the gate cannot say. The digest carries the savings and costs no facts — visible recall drops to 0.769 while total recall holds at 1.000, which is Tier-1 drawn to scale: a quarter of the facts moved behind a handle rather than being deleted. Lossless-only measures 0.0% here, reproducing what the live 4-arm A/B found on real traffic, and is the honest answer to "can I have the savings without the digest?". And the curve bends exactly once, at the rung that issues no handle.

The rungs are what the proxy runs, not a reconstruction of it. Every rung enforces reject-if-bigger per block and by tokens, as _apply_tier0 does — a run-collapse marker can cost more tokens than the whitespace it removes, so a curve measured without the guard would report savings the proxy would never take. It rescues 0 of 112 blocks on this corpus, which is why it also gets a direct unit test rather than relying on the corpus to contain an inflating block. The rung once labelled byte-exact is renamed tier-0 only (what it really was), and a subscription rung added that calls the adapter's own verbatim branch — the live verbatim path also applies the in-context structured folds for older non-recent tool output, and those are most of what a subscription user actually saves. It still measures 0.0% here because this corpus is prose logs rather than tabular output, and that is now a real measurement instead of a reconstruction that happened to agree.

Reach: AutoGen, ASGI, and two more agent presets

Two new integrations, both duck-typed and dependency-free like the rest of the package: distil.integrations.autogen (compressing_tool, DistilModelClient, compress_messages for Microsoft AutoGen) and distil.integrations.asgi (DistilMiddleware, for any app that hosts its own LLM-facing endpoint on Starlette, FastAPI, or another ASGI 3 framework). compressing_tool's wrapper carries functools.wraps plus an explicit __signature__, so FunctionTool's signature/annotation-based schema generation sees the original function, not *args, **kwargs. DistilMiddleware enforces its body-size cap while reading rather than after buffering the whole oversized body, and on the cap, a client disconnect mid-body, or a fail-open (malformed JSON, unrecognized shape, compression error), replays the exact original event sequence and headers instead of collapsing the stream into one synthesized chunk.

wrap gained presets for GitHub Copilot CLI (copilot) and Kimi CLI (kimi), and goose now also gets ANTHROPIC_HOST alongside OPENAI_HOST — AGENT_PRESETS supports multiple environment variables per preset for this.

The site: a changelog page, a claims ledger, and a withdrawn number

CHANGELOG.md only ever existed in the repo; the site had no changelog page at all. scripts/build_changelog_page.py renders docs/changelog.html from it, verified byte-for-byte by a test so the two cannot drift. The nav was hand-duplicated on every page and had drifted accordingly — benchmarks.html was a near-orphan reachable only from the FAQ, the Learn course was two clicks deep, and nothing on-site pointed at SECURITY.md, the security whitepaper, or the deploy-security guide. scripts/site_nav.py now renders the canonical topbar and sidebar from one data structure, with a checker and a test asserting every page matches; a new docs/security.html points at the three source-of-truth documents rather than restating them. The search index is rebuilt across 38 pages and 761 headings, and the site serves light mode with JavaScript disabled (the theme was set entirely by an inline head script, so a light-preferring visitor with no JS got a dark page).

No number without an entry. docs/claims.json records one entry per reader-facing number: which page, a stable anchor or verbatim snippet locator rather than a line number, and a status — verified, stale, wrong, or unsourced. The first pass found 21, of which 6 were verified, 2 stale, 1 wrong and 11 unsourced. tests/test_site_claims.py freezes the count, so a claim cannot be added or removed without a reviewed diff, and re-checks every locator against the live page. The count moves with each release that publishes a number; the pages added here brought their own entries with them, including the curve's provenance line, so a regenerated curve fails the test until the page is updated.

The one entry marked wrong was distil's own: the "1.13.0, 116 sampled requests, A/A 31/31, 100%" shadow figure on the getting-started and FAQ pages and in README.md. 1.51.1 traced that sample to replays that were silently failing on signed thinking blocks, so it is withdrawn and replaced with the current reading — 44 A/B and 11 A/A samples, 81.8% raw against 84.8% self-agreement, below the 50/30 reporting floor and explicitly not a verdict — backed by a content-free shadow.jsonl summary. The head-to-head and coding-agent benchmark tables are re-run dated 2026-09-04 against current competitor versions rather than June's, with the raw outputs and versions committed under benchmarks/results/2026-09-04/; the June table is kept as a dated historical block rather than deleted. distil's 0.0% digest row there is linked to the exact-quote guarantee that causes it. The site also stops claiming the shadow replay pins temperature 0, which it does not on real traffic.

The paper: the experiment we promised, and the scope we omitted

§5 has promised E3 (leave-one-domain-out distribution shift) since the first revision and §6 never had it. The stated blocker was real but applied only to the per-turn unit, where tau-bench and SWE-bench were graded by different models. The E8 trajectory outcomes have no such problem — one deterministic official SWE-bench harness, and all 500 instance ids carry a real domain label — so E3 runs at the trajectory level, on the E10 certificate, offline, for $0.

The result is a null with known power, and the control is what makes it readable. The bound holds on 8 of 12 repositories (66.7%), far under the 95% target — but the same certify-then-check procedure on random same-sized blocks attains only 72.7%. It fails about as often with no shift present. A permutation test on between-repository dispersion finds no heterogeneity (p=0.43 divergence, p=0.74 harm, 2000 reps), and no repository is individually significant (smallest of twelve p-values 0.07, uncorrected). The apparent failure is a mismatch of units: a (1−δ) bound constrains a population risk, and a 22-instance empirical rate carries a ~7pp standard error.

Scope honesty in the same pass. The abstract now says the 15.7% @ α=0.15 headline is SWE-bench edit-localization, not agent workloads in general. E1 gains a Wilson-CI table whose aggressive rungs' intervals overlap almost completely, so the figure's ordering is suggestive rather than established, and the caption says so. A new experiment-provenance table gives date, distil version, model, n and cost for every experiment: E1–E14 ran on 0.24.0–1.8.0.dev0 in June and July 2026, and 1.45.0, 1.50.x and 1.51.0 have changed compression since, with the expected direction of effect stated as reasoning rather than measurement. E7's cost is corrected from $50.04 to $67.31 — the old figure omitted condition E entirely. LLMLingua-2's absence from E1 is explained rather than left to inference, and the E8 artifacts' missing Headroom package version is flagged as a gap rather than guessed at. Related work gains the 2026 cluster with arXiv ids, and provider-native compaction as a first-party baseline.

A new section names the three measurements that revision does not have — degradation curve, adversarial robustness, output-token accounting — each rendering [pending 1.52.0], so a half-finished refresh can never read as a result. This release produces all three (bench --curve, validate --adversarial, and the shadow paired runs); folding them back into the paper is not done here, and those markers still render pending.

CI is now the canonical deterministic paper build: paper-build uploads both tracked PDFs at the same SOURCE_DATE_EPOCH pin the staleness check rebuilds with, so an author with no local TeX Live can satisfy the gate by downloading the artifact and committing it. Determinism was verified rather than assumed — two runs over identical paper source produced byte-identical main.pdf.

Also

distil onboard hands the terminal to distil wrap -- claude with a plain subprocess.run, and the terminal delivers SIGINT to the whole foreground group. Claude Code uses the first Ctrl+C to cancel a turn rather than to exit, so the same key raised KeyboardInterrupt in onboard's parent and dumped a traceback over the still-running agent. onboard now installs the same no-op Python-level SIGINT handler wrap's parent already uses, and restores the caller's handler after the handover.

1.51.1 — shadow mode was measuring the wrong subset of your traffic

Every shadow replay carrying a prior-turn thinking block failed with 400 Invalid signature in thinking block. A thinking block is signed and bound to the request that produced it, so replaying that conversation as a new request makes the provider re-validate a signature that cannot be valid. Verified live, including the workaround that does not work: the signature is checked even when the replay itself sets thinking off.

Claude Code runs extended thinking by default, so most turns of a real session carry these blocks. Shadow could therefore only sample the minority of turns that had none — which did not merely shrink the sample, it biased it. Every decision-equivalence rate distil has published from live traffic was computed over whichever turns happened to be thinking-free, and the A/A control never reached its own n=30 reporting floor. Measured before this fix: 594 of 1,571 replays failed, ~94% of them 400s.

Replays now strip thinking/redacted_thinking blocks from both sides identically. That is sound for the question shadow asks — "did compression change the agent's next action" — because prior thinking is regenerated by the model rather than read back as evidence, and neither side is advantaged. An assistant turn left with no content is dropped rather than sent empty, which the API rejects.

SIG_VERSION goes to 4, per the rule stated on the constant itself: bump on any change to how a signature is computed or how the compared sample is generated. v4 rows cover traffic v3 could never reach, so the two are not comparable and v3 rows are discarded rather than pooled. Your existing shadow history will therefore reset — that is the point: it was measuring a biased subset.

Also: replay_failed and signature_none_skipped counted the same event twice (594 vs 592, identical in every version — the shape of a double-count, not two causes). A skip is now recorded only when the replay actually succeeded and carried no extractable decision.

1.51.0 — the shape every fold missed

Promotes the 1.50.x line to GA and adds one new compression path. The correctness work (1.49.0–1.50.2) soaked as v1.50.1rc1/v1.50.2rc2; the embedded-JSON fold below landed after rc2 and so ships on its tests and the certificate gate rather than on soak time.

Embedded JSON

distil's columnar folds required the whole block to be a JSON array. Real tool output rarely is: a gh api dump, an MCP result, a curl | jq tail and most CLI wrappers surround the payload with a status line, prose, or a trailing summary.

Measured before this existed: a 60-record array wrapped in two sentences compressed 0.0%, while the identical array alone compressed ~70%. The machinery was already there — nothing ever handed it the span.

fold_embedded finds top-level balanced [...] spans with a quote- and escape-aware scanner (JSON nesting is not regular; a greedy regex would swallow everything between the first [ and the last ]) and folds each one in place, leaving the surrounding text byte-for-byte intact — so the prose an agent reads for context survives.

json-in-prose   2405 -> 507 tok   78.9% saved   (was 0.0%)
gh api dump     2083 -> 371 tok   82.2% saved   (was 0.0%)

Wired into both fold chains. The adapter keeps its own, so a transform added only to tier1 compresses in tests and 0% in production — the exact split that once left HTML at 0.0% on real traffic. A test pins the adapter path specifically.

It runs last: the whole-block folds are strictly better where they apply, and a block that is an array must not be encoded twice in two different markings. Reject-if- bigger and the DECISION: carve-out are honoured like every other fold, and the certificate gate still passes non-inferior on every trajectory.

1.50.2 — a counter could end a session

Session survival, not savings. A proxy sits in the request path of a live agent session, so the contract is narrow: distil may fail to compress, but it must never fail to serve. It could.

Token accounting — _count_messages, whose entire output is a response header and a ledger row — ran unguarded in the request path. Compression was protected, the retention meter was protected, but the counting was not. An exception there escaped the handler and closed the connection: the client saw RemoteDisconnected and the turn was lost. A tokenizer edge case on unusual content would have ended a live session over a number nobody reads in the moment.

Now guarded, and the failure is honest about itself: a request whose counts did not compute goes unbooked rather than entering the savings ledger with a fabricated zero. A wrong number in the savings history is worse than a missing one — every published percentage derives from that ledger.

Three end-to-end tests drive the real handler over a real socket, because a try/except in the source is a claim and only a served response is a measurement: a crash inside compression, a crash inside accounting, and a body distil cannot parse (forwarded as-is, so the provider's own error reaches the agent rather than one distil invented).

A failure that healed itself was still reported as a failure

distil recorded 5,397 non-2xx requests across 18,455 and discarded the reason every time. Only the status survived, so the one question a broken session raises — why? — had no answer in distil's own logs, and dissect guessed on the user's behalf: "upstream errors or rate limiting."

That guess spans two opposite diagnoses. "Your conversation outgrew the context window" means nothing is wrong with distil; "distil sent a malformed body" means everything is. A tool whose pitch is proof rather than trust should not be unable to tell those apart, and this one wasn't: the missing field caused a real misdiagnosis of a live session, where a size correlation looked causal and was not.

Two things now get recorded:

  • The provider's error type (invalid_request_error, rate_limit_error, …) — a short enum. The human message is deliberately NOT stored: it quotes request content ("prompt is too long: 231721 tokens > 200000"). Streaming responses are relayed frame-by-frame and never buffered, so those record http_<status>.
  • Whether a retry recovered it. A non-2xx immediately followed by a same-size 2xx is the SDK doing its job. On the session that prompted this, 100% of the 400s were retried and succeeded — a ~30% "failure rate" in which nothing was ever lost. Reporting that as failure sends people hunting a bug that already healed.

Sessions recorded before this release report no reasons rather than an invented one.

1.50.1 — the tokens distil could not see

Extended thinking carries its payload under thinking, not text, so distil counted it as zero — a 1,000-character block scored 0 tokens. On Claude 4.6+ prior-turn thinking is re-sent as input and billed every turn, which means real billed tokens sat in neither the before nor the after side of every percentage distil reports.

The blocks are still never rewritten, and that is not timidity: the provider pins each block by its signature and re-expands the original server-side, so editing the text achieves nothing and risks the signature being rejected on replay. But a cost distil cannot reduce is exactly the cost it should not hide. Thinking now appears in the token baseline and in the eligibility census as thinking_billed.

1.50.0 — the model that could never learn, and the store nobody could see

Two capabilities that were present in the code and unreachable in practice.

Query-aware salience can finally train. The shipped query_weights.json never existed on any install, and the reason was structural rather than a cold start: label collection ran only under --expand, and the labels themselves were expand events. So the flywheel could only learn from the one configuration an independent benchmark tells users not to adopt — and never at all on a subscription, where the digest tier is off and no expand can occur.

Shadow mode already produces a better label. An A/B verdict says whether compressing this request changed the agent's next action — the property the certificate is about, not a proxy for it — and shadow runs by default at 2% on every configuration. Training now treats a shadow decision-change as a positive alongside an expand, joined on the request digest both sides already compute. A/A rows are excluded: those re-run the same compressed request twice, so a disagreement there is provider nondeterminism, and learning from it would teach the model to keep content because the sampler was noisy.

distil memory — the cross-agent recall store, made visible. A handle minted while compressing for Claude Code was always expandable from Codex, Gemini, or any MCP client reading the same DISTIL_HOME: the store is machine-wide, encrypted at rest, and needs no database. Nothing surfaced it, so nobody knew, and nobody could tell when it was full. distil memory reports what is stored, how old it is, and warns within 10% of the cap — the point where the oldest handles start being evicted and a stub in an agent's context quietly stops being expandable. --clear empties it.

Deliberately NOT built: an embedding-backed semantic memory. The comparable feature elsewhere costs a vector database, a graph store and a sentence-transformers model. distil ships zero runtime dependencies and installs anywhere on day one, and the recall property users actually need — get me back the exact bytes that were folded — is served by a keyed store rather than a similarity search over paraphrases.

1.49.0 — the agent said "done" and wrote nothing

Fix this one. An independent benchmark ran 75 agent sessions (Claude Code, Opus 5, ground truth taken from the API's own usage fields rather than any tool's dashboard) and found distil wrap --expand completing 6 of 15 coding tasks against bare Claude Code's 13 — seven runs ended with an empty diff: the agent ran 18–51 turns, wrote nothing to disk, and reported success. Zero empty diffs occurred across the other 60 runs. On the long refactor it failed all three times, and the repository's 1,371 existing tests still passed, because nothing had been modified.

Four defects, all now fixed and each pinned by a regression test verified to fail without its fix.

The agent never ran its own tool call. Claude Code emits parallel tool calls, so one assistant message routinely carries both an Edit and a distil_expand. The expand loop assumed distil_expand was the only tool call in the turn and spliced it: the replay answered only the expand block — leaving the client's tool_use unanswered upstream, an API-contract violation — and the continuation's stop_reason: end_turn replaced the turn's real tool_use. Claude Code executes tools only on stop_reason == "tool_use", so it delivered the Edit, ran nothing, printed the continuation's "all done", and stopped. A turn carrying a client tool call is now terminal: relayed verbatim, stop_reason intact, no re-query. Guarded on all four paths — streaming, Messages, Responses, Gemini.

A failed recovery truncated the stream. When the re-query failed, the connection closed with no message_delta/message_stop. A terminator-less SSE message is truncated, not finished, so the SDK retries it invisibly — wall-clock burned with no progress and no error the user can see (the benchmark's 808-second, zero-write signature). Every exit now emits a terminator.

A file read could be digested before it was edited. Recency is positional, so a read from three turns ago was digestible — but the agent must still reproduce that text character-for-character in an Edit(old_string=…). Once the read is a digest no exact match exists, so the edit silently fails or is never attempted. The exemption is now keyed on provenance — results answering Read/Grep/Glob and their MCP equivalents stay byte-exact at any age. Logs and test output are unaffected and still digest.

Compression that saved nothing still cost cache. On a subscription (lossless-only by design, 0.0% savings correctly reported) distil still wrote 1.56× baseline cache-creation tokens, 2.52× on a short session. Two causes, both fixed: the distil_expand tool was injected only once a handle existed, so the tools array — which Anthropic caches ahead of the system prompt — changed shape the turn compression first fired; and an unmodified body was re-serialized rather than forwarded byte-for-byte. Injection is now session-sticky, and unchanged bodies forward their original bytes.

distil cache also stops sending users to the wrong place: it blamed prefix drift on "a tool list whose order varies" upstream, while distil's own tool list was the one varying. It now names distil's own causes first.

Hook mode can finally account for itself. The same benchmark found hook mode the strongest distil arm — 12/15 tasks, 0.94× baseline cache writes, the highest cache-hit rate measured — and the only one that wrote no ledger at all, so its effect was visible only in the provider's billing. It now appends a content-free receipt per compressed result, readable with distil hook --stats.

  • distil wrap -- cursor (and Copilot, Cline, Continue, Windsurf, Zed) now says these are editor extensions with no process to wrap and no base-URL variable they read, and names the proxy path that does work — instead of setting a variable nothing reads and reporting success.
  • A missed expand handle is no longer logged as a successful recovery. The placeholder was being passed to the learning signal as content, so a failed recovery trained the keep-model that the block it could not return was safe to drop. distil dissect gains expand_missed, so expand_resolved finally has a denominator.

1.48.1 — a page that answers "which one am I?" before the jargon does

Docs and one nudge. Nothing new is callable and no behaviour changed, so this is a patch: if you are tracking releases for capability, skip it.

distil has seven ways to run and the names are all jargon — wrap, hook, proxy, MCP, library, gateway, statusline. A newcomer had to learn our vocabulary before they could tell which one applied to them. docs/which-mode.html inverts that: two plain questions (how do you pay, what are you running) and the answer falls out, with an animated decision tree and honest savings ranges per surface.

The page leads with the thing most compression tools bury: the mode barely matters, what your tools print matters enormously — JSON 28–33%, repeated log lines up to 99%, prose ~0%, and 0.00% on distil's own corpus.

  • distil onboard now closes with a pointer to that page. It already detects your OS, agents and billing, but it could never mention MCP, the library, or the gateway, because those cannot be detected — which is how the MCP server (the only thing that works in Claude Desktop) stayed invisible to most installs.
  • The provider-compaction paper gains an affiliation and its arXiv category (cs.SE primary, cs.LG cross-list), plus the submission step that is easy to get wrong: upload the .tex and its generated macro fragment, or every number renders as --.
  • A nav-markup guard, after review caught two links sharing one <li> — a defect that had already shipped across 35 pages in the previous release's nav.

1.48.0 — the saving on a subscription is the window, not the bill

A reader was right and our copy was wrong. distil's README said a flat-rate Pro/Max plan gets "context and latency, not the bill". But fewer tokens per turn means more turns before the rate-limit window closes, and on a flat-rate plan that window is the currency. The copy is fixed and this release makes the saving real.

The proxy can't help there, by design. Anthropic's consumer terms (§3, item 7) restrict automated access on subscription credentials, so distil deliberately runs --lossless-only on a subscription and measures 0.27%. Your account isn't worth a few percent.

So use the door that's open. distil hook --install wires a Claude Code PostToolUse hook: the agent compresses its own tool output, in its own process, through a documented first-party extension point. No proxy, no credentials touched. Because a hook sees each result once and cannot rewrite history, compression is append-only by construction — the fix direction our own cache-busting investigation identified, enforced by the platform rather than by our discipline.

Measured on a paired live A/B, both arms answering correctly: tool_result −38.6%, cache_creation −67.4%, cost-weighted −68.3%, decision-equivalence 5/5. Critically cache_read did not collapse — the proxy digest's failure mode does not reproduce here.

And the number that doesn't flatter us: on distil's own eval corpus the hook saves 0.00%. Savings are shape-dependent — verbose JSON 28–33%, duplicated log runs up to 99%, prose and unique-line output ~0% — and our corpus contains none of the winning shapes. Published because quoting only the favourable fixtures is the overclaim we criticise in others.

  • distil hook --install / --selftest / --uninstall — idempotent, preserves foreign hooks, refuses to clobber an unreadable settings.json
  • distil quota — the rate-limit windows, read-only, fails open to "unavailable" rather than a fabricated zero
  • docs/subscription.html with an animated diagram and a troubleshooting table
  • The provider-compaction paper (docs/paper/provider_compaction.tex), arXiv-ready, every number generated from run artifacts by a script that refuses to emit LaTeX unless each report's protocol hash matches its pre-registration

Credit where it's due: Headroom shipped subscription quota telemetry against this endpoint before we did. They built the instrument for the subscription user's real currency while our copy still said that currency didn't count.

Windows note: three defects in this work were Windows-only (os.uname() in the code, then in the test that proved the fix, then pytest IDs blowing the 32,767-character environment-variable cap). All fixed, and the last one is now fenced by a guard verified to actually fail.

1.47.0 — the key file was world-readable, and the audit trail wasn't watching

Two things an enterprise security review would have caught before you did. The master key that decrypts every restore blob was written at the process umask — measured 0o644 — and only chmod'd to 0600 afterwards. Any local user reading it in that window decrypts everything. The gateway key file had the same shape: a polling thread caught its temp file at 0o644 during a 400-key issue loop. Both now create 0600 at open time, so the window does not exist. Re-probed after the fix: 1,855 samples, only 0o600 ever observed.

The gateway had no audit log at all — grep -c audit over gateway.py and authz.py returned 0 and 0. There was no record of who authenticated, which tenant used which key, or what was refused. That is the first thing a security questionnaire asks about a shared gateway, and SOC 2 CC7.2 / ISO 27001 A.12.4 both require it. There is now an append-only, flock-guarded, 0600 JSONL trail of auth.ok / auth.fail / rate.limited / key.issued / key.revoked, read with distil gateway audit (--json for SIEM ingestion).

It covers every refusal path, not just the convenient ones: no credential presented, no key store configured, invalid or revoked key, OIDC token rejection, OIDC RBAC denial, OIDC success, both RPM paths, and both daily-token quota paths. The first cut missed most of those — a deployment on OIDC got an empty audit log no matter how much traffic it refused — and that gap was caught in review before shipping.

Content-free by construction: identifiers and outcomes, never prompt text, completion text, tool output, or the raw dsk- token. Writes fail open, because bookkeeping must never drop a customer's request.

Keys can now expire. distil gateway keys issue --tenant acme --expires-in-days 90 gives a key a bounded lifetime, enforced through the same is_active chokepoint that enforces revocation — so no code path can honour an expired key by checking only revoked. Keys issued without an expiry still never expire and pre-expiry key files load untouched, so nothing changes for existing deployments. keys list reports expired separately from revoked: an operator debugging a sudden 401 should not go hunting for a revocation that never happened.

Compression got faster without changing what it produces. _xor_stream XOR'd byte-by-byte in a Python loop; must_keep ran three to four regexes on every line, where one pattern cost more than all the others combined. Vectorising the first and putting literal prefilters in front of the second takes a 3.2 MB context from 1014 ms to 719 ms per request, with savings percentages unchanged (95.8 / 92.3 / 81.9 / 48.3 at 60 / 30 / 12 / 4 turns). The encryption output is byte-identical — verified on 627 cases including leading-zero and partial-block torture — and the keep decisions are identical across a 400,000-case differential test.

Minified JSON was leaving most of its savings on the floor. A single-line payload returned before it ever reached the columnar folder that already existed for it: every REST tool call, every curl | jq -c. Minified JSON goes from 10.5% to 57.5%, nested records to 91.2%. The fold stays in-context lossless — all 2,500 field values across 500 records appear literally in the output — declines when values contain tabs or newlines so no column can shift, and leaves the recency carve-out byte-identical.

Also fixed: os.write is permitted to write fewer bytes than it is handed, and the first cut of the key-permission fix discarded that return value. A short write left a truncated key, the next load minted a fresh one, and everything encrypted with the first became unreadable. Path.write_bytes loops; the hand-rolled version did not.

Known limitation, documented in atrest.py where someone will hit it: concurrent first touch of a brand-new key store can still leave two creators on different keys. Five attempts to close it each passed on Linux and macOS and broke Windows a new way; it wants a Windows runner in the loop rather than a sixth blind fix.

1.46.0 — every managed install was running blind

If you installed distil the managed way, your proxy has never sampled a single decision-equivalence check. distil wrap defaults --shadow 0.02 and --retention 0.05; distil proxy defaulted both to 0.0. The com.distil.proxy launch agent runs proxy. So the two quality loops that make distil's central claim checkable — shadow's live decision-change rate and the fact-retention meter — were off on every managed install since the daemon shipped. A 93-minute, 767-request session sampled 0. The same work under wrap would have sampled ~15.

Both defaults now match wrap, so existing installs are fixed by upgrading; no reinstall. --shadow 0 / --retention 0 still opt out. Retention costs nothing either way (in-process scan, counts only); shadow spends ~2% extra tokens on sampled requests, which is the price of having evidence at all.

distil dissect reported three different denominators for the same session

Reading one report end to end, four numbers disagreed:

  • The request-detail lines summed all requests while the savings headline counts booked (2xx, non-retry) ones. 153 unbooked retries contributed 9.85M overhead tokens and 3.75M of "savings" that were never billed.
  • overhead_share divided by the pre-compression payload — measuring the fixed tax against tokens distil had already removed. It read 28% where the session's own numbers implied 42%.

Both now use the booked population and the post-compression denominator. "Savings by mechanism" reconciles with the headline it decomposes: a 3.75M discrepancy became 72k (0.34%), which is heuristic-vs-ledger tokenizer rounding. Same defect shape as 1.44.0's, which fixed it in savings and did not carry dissect along.

"Everything summarized stays recoverable" was false while it printed

The restore store has a 500-blob LRU cap, not just a TTL. A single session folded 704 blocks and evicted 204 of them mid-session — including its most re-used fold, one referenced by 272 requests — while the report promised full recoverability. It now reports the measured count. The cap is configurable (DISTIL_RESTORE_CAP, default raised 500 → 5000). A cap of 0 now means "no count cap" instead of evicting everything, which is what [:-0] did.

A flat 0.0% per-model row is now explained, not left looking broken

One model showed exactly 0.0% across 152 requests. That is policy, not a failure: the traffic was 100% user-role string content with no tools — subagent calls — which routes to Tier-0 lossless by design. The row now says so.

Papers and published claims, audited against the code

  • Both papers compile with zero undefined references. Fixed a missing \bibitem, a dangling cross-reference, and a figure quoting a number no artifact produces.
  • The headline macro fallback silently rendered a different operating point as the headline (0.1% certified savings) when the generated macros were absent. It now warns at build time.
  • docs/paper/generated/headtohead_orig.tex was untracked, so the built PDF rendered numbers absent from version control.
  • The paper described cache-monotonicity as holding by construction. That property was false in shipped code until 1.45.0 fixed it. The passage now describes the breakpoint anchoring that actually ships, and reports the measured failure.
  • The NeurIPS variant was missing the live-validation negative result — the section where a real grader failed the live margin twice, reproducibly. Ported.
  • The security whitepaper said OIDC and role-based access control were "not implemented"; distil/authz.py implements both. Corrected to the honest gap (SAML, SCIM).
  • Docs said the proxy needs Python 3.11+; the floor is 3.9.
  • The site advertised 25.3% aggregate corpus savings over "8 domains" (listing 7). distil bench reports 48.4% over 9 domains. We were underselling by ~22pp against our own free, offline gate.

1.45.1 — the explainer command was the one reporting uncalibrated numbers

distil dissect compared the raw heuristic token estimate against billed usage, so its "off by >50%" tripwire fired on essentially every calibrated session. Calibration is applied on every other surface; dissect is the surface whose entire job is explaining the numbers, and it was the one quoting them uncalibrated.

1.45.0 — the compressor was rewriting the cache it was supposed to protect

On cached agent traffic, distil cost about twice what sending nothing compressed would have. Provider prompt caching discounts a repeated prefix by ~10x, but only while it is byte-identical. distil kept the freshest tool outputs verbatim so an agent could see its latest observations exactly — using a window counted back from the end of the message list. That window slides. Every block was protected while fresh and digested one turn later, which rewrote a message the client had by then committed to its cached prefix. One changed block, and the whole prefix is re-billed at the 1.25x write rate instead of re-read at 0.1x.

Measured against the live API, 10 turns, per-arm and per-run salt so no arm reads another's cache entry:

arm                     cache_create   cache_read        $   vs control
control (no proxy)            29,446      146,572   0.0517
lossless_only                 29,450      146,608   0.0517       +0.0%
digest                        85,876            0   0.1076     +108.1%

Zero cache reads, on every turn. Compression halved token volume and still doubled the bill. --session-delta did not help. This is the exact failure mode distil has always attributed to naive compressors.

docs/CACHE.md claimed the property held "by construction", and two shipped CLI diagnostics told anyone debugging drift that the cause must be upstream in their own prompt assembly. Those are retracted.

The rule now is exempt only content the provider will not have cached. Anthropic caches what the client marks, so recency is anchored to the last cache_control breakpoint — with no marker there is no prefix to invalidate and the previous last-k window is kept. OpenAI and Gemini cache implicitly and commit everything they are sent, so they digest on first sight instead. There were four copies of the sliding window; all four now route through one rule.

A second cause was found by looking for other per-request inputs: query-aware salience was re-deciding how already-cached blocks compress, because intent terms come from the newest user turn. Asking a different follow-up over the same history rewrote cached message #0 with no recency window involved. Intent is now scoped to uncached content — a block's first rendering still gets it, later requests reuse those bytes. On OpenAI and Gemini that means query-aware salience no longer applies at all; those providers cache everything they receive, so there is no uncached region for it to work in. That is a real reduction in what the 1.16.0 feature does, taken deliberately against a prefix rewrite on every question.

After both fixes, same harness: -60.3% vs control, cache reads intact.

Pinned by append-only invariant tests on all three adapters — compress a growing conversation and assert no earlier turn's bytes ever change. The existing gates could not have caught this: reversibility and decision-equivalence are properties of a single request, and this was a property of the sequence.

Digesting the freshest observation is what the certified strategy already does, so serving moves toward certification rather than away from it. It is still a live behaviour change for OpenAI and Gemini users and deserves a shadow run before it is treated as settled.

Also in this release: lossless_only measured +0.0% against no proxy at all — it forces Tier-0-only, and serve()'s docstring claimed the opposite, which hid why live savings were flat. Corrected. --expand lifts the force.

1.44.0 — the savings number could not explain itself

A low percentage was indistinguishable from a broken compressor. distil stats would report 0.4% and stop there, and from the outside that reading has four completely different causes: a request that was mostly the model's own prose (which distil never rewrites, by policy), mostly recency-exempt tool output (a carve-out that shrinks as a conversation grows), mostly fragments below the digest threshold, or a digester that was handed real work and declined it. Three of those are the design holding. One is a defect. The saved-token count is identical in all four, so the only way to tell them apart was to reason from the outside about what the traffic probably contained — and that reasoning is wrong often enough to send an investigation in the wrong direction for hours.

Every block is now attributed to the gate that actually claimed it, and the counts ride on the per-request record. distil dissect sums them and prints the answer next to the number that prompted the question:

why: the model's own words (never rewritten) 54%, digester declined 25%,
     digested 12%, freshest tool output (kept byte-exact) 9%

A declined share of 5% or more is called out as the compressor's to explain. Otherwise a protected majority is named as the design holding rather than a fault. Declined wins the tie deliberately: a session can be 63% policy-protected and refusing a quarter of what it was handed, and summarising that as "mostly protected, all is well" would file a real defect under an exoneration.

The counts are taken at the decision points themselves, not by a second pass that re-derives the rules — a parallel implementation is free to drift from the compressor it describes, and a report that quietly disagrees with behaviour is worse than no report. The first draft of this predicted the branch order and got it wrong: HTML extraction runs before the minimum-lines gate, because minified markup is one enormous line, so every browser fetch would have been filed as "too short to digest".

Sessions recorded before this ships render no why: line at all, rather than showing zeros. "No data" and "nothing was eligible" are different claims and only one of them is true.

Billed tokens never said what a request cost your plan. Provider responses carry anthropic-ratelimit-* headers — the only first-party signal for how a request draws down a quota, as opposed to how many tokens it contained. The two diverge sharply on cached traffic, where a read is billed at roughly a tenth of the metered price, and whether a subscription cap discounts it the same way is not published anywhere. Those headers are now recorded per request, including on the 429 that carries the most informative quota state of any response. Captured at all three upstream reads, because they are genuinely separate code paths and the --expand splice path is the one real agent traffic uses.

distil stats quoted a digest rate from a window it did not belong to. The rate beside the trim was computed over all recorded history while the trim itself covered the recent window, so a mode enabled today was described with a number earned by months of traffic under a different one. Both figures now come from the same window.

1.43.0 — the safe default could not say what it cost, and the reload could strand you

distil default --always-on could take macOS down. The reload boots the job out, replaces the plist, and then demanded a free port before bootstrapping. Since 1.42.0 the plist declares Sockets, so launchd owns the listener and the SIGTERMed child that inherited the descriptor can still be holding the port at exactly that moment — a held port there is the design, not a foreign process. The check fired on that ordinary case, and it fired after the bootout: the job was already unregistered, so the machine was left with no proxy and nothing to restart it, while the command printed "your existing setup is untouched". It was reproduced by nothing more exotic than switching modes. Linux got this fix for the upgrade path in 1.42.1; macOS never got the equivalent. The port settle is now a wait, not a verdict — if the port really is stolen, the bootstrap still runs and the failure is reported with the job registered, so KeepAlive keeps retrying instead of stranding the machine. And a failed reload no longer claims nothing changed; it says the previous service is stopped and gives the command back.

The status line called a fully-routed session a bypass. With an always-on install, the settings-file ANTHROPIC_BASE_URL pin outranks the env var distil wrap injects, so the agent talks to the always-on proxy — which stamps ledger rows with its own session id and cannot know the agent's. The wrap session's traffic marker therefore never flipped, and after the grace period the line read ⚠ wrapped, agent bypassing proxy for the life of a session whose traffic was flowing through distil the entire time, with nothing the user could do to clear it. The marker only ever proved "this wrap's own proxy saw nothing"; the warning claimed "no distil proxy is seeing this traffic". The evidence now matches the claim.

A subscription default that could not say what it costs. A flat-rate session runs lossless-only so no digest is left unrecoverable and nothing is injected — deliberate, and unchanged. But on the machine that prompted this it meant 8,319 runs at 0.27% while that same ledger held 2,353 digest runs at 52.27%, and all distil ever said was that near-zero "usually means lossless-only". No figure, no sample size. That line printed on every one of those runs and was read past every time. It now quotes the rate you earned on your own traffic, and states both halves of the trade — the safe default does not touch the request, and --expand does — because you cannot weigh a choice shown only one side of.

You could opt into tool injection with no notice. The disclosure was gated on the --lossless-only --expand spelling only. The route the docs actually hand a subscription user, distil default --mode expand, leaves that flag False, so the notice never fired — and injection into a first-party session is precisely what the safe default promises not to do. It is now keyed on real billing and fires on every route into it.

policy.may_inject_tools() is gone. It stated that injection is PAYG-only, a rule distil does not implement, and nothing ever called it while a passing test asserted it — so the drift between what distil said and what distil did was invisible in CI. An unenforced policy with a green test is worse than no policy. ADR 0006 records what the boundary actually is: consent, not recoverability. distil does not modify a first-party session unasked. Recoverability is still why a digest is safe once opted into; it is not why the default exists.

Two more agents, and one needed more than a preset. distil wrap routes Grok CLI (GROK_BASE_URL; the /v1 belongs to the base URL there, unlike the OpenAI SDK which appends it) and OpenHands. OpenHands reads LLM_BASE_URL only when run with --override-with-envs — without it the agent reads its own settings file and ignores the environment entirely, so a preset alone would print success while routing nothing. wrap checks argv and says so before the session starts. The IDE extensions (cursor, copilot, cline, continue, windsurf) remain deliberately absent — no argv to wrap and no published env contract — and a test now pins that.

Oversized images can be downscaled, behind a recovery handle. Duplicate elision cannot touch a first occurrence, so a 2400x1200 screenshot still cost its full ~1,600 input tokens every turn. ADR 0007 permits downscaling in exactly one form: the original goes into the RestoreStore first, and the smaller image is emitted as a pair with a note stating what changed and the handle that returns the untouched original. A downscaled image emitted alone would be silent loss and is not allowed. Ships off, behind its own certificate, with no bundled certificate — vision's shipped result certifies a provably identical image, while this is a claim about whether losing detail changes decisions on your screenshots. Needs the optional [image] extra; inert without it.

Networks that inspect TLS no longer make distil look broken. Behind a corporate MITM appliance every outbound connection is re-signed by an internal CA. curl and the browser read the OS trust store; Python does not — so distil failed certificate verification while every other tool on the machine worked. The upstream opener now honours DISTIL_CA_BUNDLE / REQUESTS_CA_BUNDLE / CURL_CA_BUNDLE / SSL_CERT_FILE, a verification failure returns the sentence that fixes it rather than a raw OpenSSL string, and distil doctor names the bundle actually in effect. A path that does not exist is ignored rather than fatal. There is no option to disable verification. See docs/ENTERPRISE.md §3b.

distil dissect was reporting almost nothing on always-on installs, and the cause was one unexported variable. The savings recorder mints a session id when distil wrap did not set one and put it on ledger rows — but every session-scoped path resolves that id from the environment, and an always-on proxy runs under launchd or systemd, neither of which inherits a shell environment. So per-request records, the session manifest and the liveness marker were all silently dropped while the ledger looked healthy. dissect reported blocks=0 on sessions worth over a million input tokens: the analysis was never broken, its input was never recorded. This was the third surface of that one cause — the status line calling fully-routed sessions "bypassing", and ledger rows attributed to the proxy rather than the wrap, were the first two.

distil learn --write puts what a session measured where the agent will read it. That analysis has always lived in a report a human reads once; the agent that produced the cost never saw it, so it repeated itself next session. This writes a managed block into CLAUDE.local.md keyed to your project: which content type dominated, how much of the bulky content it was, and the move that reduces it. It writes only to a *.local.md file unless forced (a tracked CLAUDE.md is reviewed by your teammates), preserves everything outside its markers byte-for-byte — including line endings and non-UTF-8 bytes — and declines entirely when no single content type reached half the measured volume. A file of plausible advice is one the agent learns to skim.

1.42.1 — the upgrade could not claim the port it had just been given

distil default --always-on was broken on Linux in 1.42.0, in both directions, and macOS was unaffected — which is why it shipped.

Upgrading from 1.41.x: before 1.42 the systemd service self-bound the port and there was no socket unit. On upgrade that old service is still running and still owns 127.0.0.1:PORT, so enable --now distil-proxy.socket cannot bind and fails — after the unit files have already been overwritten. The user is left with new units, an error, and the old proxy still holding the port.

Re-running on an install that already works: the socket unit holds the port by design and keeps holding it after the service stops — that is the entire feature. A reload that stops only the service and then demands a free port therefore aborts every ordinary re-run with "something else is listening", pointing the user at a culprit that is distil itself.

Both are one fault: ordering. The reload now stops both units, waits for the port, starts the socket so systemd owns the listener, then starts the service, which receives the descriptor instead of binding for itself. No single enable --now can express that sequence, which is why the command string service_spec used to return could never have been right — it is now a marker, and the last copy of the racy launchctl unload; launchctl load is gone from the codebase.

Three more faults in the same area:

  • Extra descriptors were leaked and hung. A dual-stack socket unit hands down more than one fd; distil wrapped the first and left the rest open. The leak was the lesser problem — connections arriving on an unaccepted listener queue in the kernel forever, so a client hangs instead of failing. A hang is worse than a refusal: nothing surfaces an error to act on. Extras are now closed on both platforms.
  • service_reload could hand the user a traceback. Its sibling guarded its subprocess calls; it did not, so a launchctl/systemctl that hung past its timeout printed a stack trace where a sentence belonged.
  • Socket cleanup was gated on the .service file existing. A socket unit can outlive its service — a partial uninstall, a hand-deleted unit, an interrupted upgrade — and it is the unit that owns the port. --undo and offboard now clean up when either is present.

Also: libc.free is given explicit argtypes. ctypes' default conversion happens to pass a 64-bit pointer correctly; "happens to" is not a contract.

Every fix has a regression test verified to fail without it. One of those tests is a correction in itself: the first version stubbed _port_free to True, mocking away the exact check under test, which is how the re-run break got past a green suite. It now drives the real function against a genuinely held socket.

1.42.0 — the outage was the test suite

distil's own test suite tore down the developer's live proxy, and that is what took a machine's API traffic out roughly fifteen times in one evening. cmd_offboard and distil default --undo call service_spec(8788, ...), which returns the real ~/Library/LaunchAgents/com.distil.proxy.plist — then launchctl unload it and unlink() it. Five tests never patched service_spec. Every full-suite run silently uninstalled the always-on service while ANTHROPIC_BASE_URL stayed pinned at the now-dead port, so every session afterwards failed with ConnectionRefused, an error that names the provider.

It was hard to see because the job came back unregistered rather than crashed: proxy.err was empty, there was no crash report, and KeepAlive cannot restart a job that is no longer loaded. Remembering to patch per-test fails open, and it failed open five times — so conftest now redirects the plist into a tmp dir and makes the destructive commands inert by construction.

launchctl unload; launchctl load was a coin flip. unload is asynchronous, so the replacement job often died on EADDRINUSE and launchd discarded it — while the legacy load still exited 0 and the orphaned old process kept answering just long enough for the routing probe to pass. distil default printed ✓ and wired the pin onto a machine with no registered job; nothing restarted it when the orphan exited. service_reload() uses the modern bootout/bootstrap API, waits for the job record to clear (the port frees the instant the process dies, while the record lingers in SIGTERMed — and bootstrapping into that returns Bootstrap failed: 5: Input/output error), retries transient EIO, and verifies a running pid. A failed start now wires nothing.

A restart is no longer a refusal. launchd and systemd own the listening socket (Sockets in the plist; a real distil-proxy.socket unit on Linux), so connections queue in the kernel backlog while the proxy restarts instead of being refused. Verified on hardware: kill -9 against the running proxy four times returned HTTP 401 every time, connect=0.26ms, zero refusals — where the same machine had previously logged a 37-second window of ECONNREFUSED.

Nine more faults in the machine-wiring path, each with a regression test verified to fail without its fix:

  • A pre-existing alias claude=... made the rc file unparseable. bash and zsh alias-expand the word before () while reading a function definition, so the managed block was a syntax error that abandoned the rest of the user's rc — PATH included. An earlier distil installed exactly such an alias.
  • os.replace destroyed a symlinked rc, detaching dotfile repos (stow, chezmoi) from the file they manage.
  • A hyphenated agent name (claude-code) is a syntax error in dash, which is /bin/sh on Debian and reads ~/.profile. A subshell capability probe now gates the definition. eval … || true does not work: eval is a POSIX special builtin, and a syntax error in one exits the shell before the || is reached.
  • The plist spliced paths and mode into XML unescaped; a & in $HOME produced a plist launchd silently refuses to load.
  • The escape hatch matched loopback by substring, so a corporate gateway at …/pools/127.0.0.1 would have been deleted and http://127.1:8788 kept.
  • The escape hatch scanned only hardcoded $HOME rc paths, missing $ZDOTDIR and any --rc path — the machines least served by the defaults.
  • remove_managed returned ok when the end marker was missing, so undo and offboard printed ✓ while the wiring stayed live.
  • _port_free probed by connecting, which fills the listener backlog: against a hung proxy it reported a held port as free, causing the exact EADDRINUSE it exists to prevent. It now probes by binding, which is what the supervisor does.
  • Nothing removed the systemd socket unit on uninstall. An orphaned enabled .socket keeps the port bound after distil is gone and keeps systemd starting a service that no longer exists.

Also: the subscription notice drops from seven lines to two (it prints on every bare wrap, and a wall of text is skipped exactly by the readers it exists for), and CITATION.cff gains the drift guard it never had — it was the one version location no test pinned, and it had already fallen behind to 1.41.1.

distil shadow-stats now reports the running build's replay-failure rate, not a lifetime average (#96). The counters accumulated for the life of the install, so failures from an already-fixed bug stayed in the displayed rate forever: one install read 42% lifetime while the rate since the fix was 5.3%. That does two things, and the second is the one that matters — a fixed bug makes the sampler look permanently broken, and the next real regression has to clear that stale noise floor before anyone notices. Counters are now bucketed by version (lifetime totals kept, as context and as a trend), and last_fail_reason becomes a fail_reasons histogram, so a mixed run says 1 failed (429×1) instead of reporting whichever failure happened to land last.

Socket activation on Linux is now proven rather than asserted. The launchd half was verified on hardware from the start; the systemd half was covered only by tests against mocked units. Two Linux-gated tests close that: an end-to-end check that does what systemd does (bind, hand the listener down as fd 3, exec, then SIGKILL the worker and demand the next request still be served), and systemd-analyze verify on the generated units — because asserting on our own generated strings cannot catch a directive that is misspelled, deprecated, or illegal in the section it was placed in. Both run on the Linux CI matrix.

A security test that had never once executed now does. test_authz.py's RS256 malformed-key assertion guards itself with importorskip("cryptography"), and the CI gate never installed it — so it skipped silently in every run since it was written. The same shape as the outage above: a green result that was green because the thing wasn't running.

What is and isn't verified, since this release changes machine wiring on two platforms. macOS is proven on hardware: kill -9 against the running proxy four times returned HTTP 401 every time, and distil default --always-on went from roughly a coin flip to 5/5 deterministic. On Linux the LISTEN_FDS handoff is proven end to end and the generated units are accepted by systemd's own parser, both on every CI run. What is still not covered by a test is systemd actually starting the units and handing the descriptor over on a live host — systemd-analyze verify proves they are well-formed, not that a running systemd wires them. Windows is unaffected by the wiring paths and is only verified not to break.

1.41.2 — the escape hatch that wasn't, and a price nobody was charged

distil default --always-on took launchctl load returning 0 as proof the proxy worked. It means the job was accepted, nothing more. On one real machine the proxy bound its port and 404'd every request — a --upstream pointed at a server that does not serve /v1/messages — and because Claude Code skips model-name validation whenever ANTHROPIC_BASE_URL is set, every session on that machine failed as "there's an issue with the selected model". The message sends you hunting the model. The fault was the base URL, and distil had written it.

probe_routing now POSTs /v1/messages and demands anything but a 404 — a 401 is proof, because it means the request reached a real messages handler. --always-on starts the service, proves the route, and only then writes the base URL; a failed preflight wires nothing and exits 1. doctor probes the same way: "the port is listening" was the old bar, and it was already true the entire time everything was broken.

The second defect was worse, because it was the way out. Both removal paths — default --undo and offboard, the command someone runs when they have had enough — looked only in ~/.claude/settings.json and matched the entry against a port: --undo against the current --port, offboard against a hardcoded :8788. Claude Code also merges project settings, and those take precedence over the home file. offboard carried a third defect of its own: it prompted only if the home file existed, so on a machine without one it asked nothing and cleaned nothing, in any file — then reported success and printed the uninstall command.

All of it was true of the same machine at once. A dead 127.0.0.1:8788 in a repo's .claude/settings.local.json survived uninstalling distil entirely, and went on killing every session started in that directory.

claude_settings_files() enumerates every file Claude Code merges, in precedence order. unwire_base_url() cleans by shape — any loopback URL is ours — while leaving a real remote gateway (LiteLLM, Bedrock, a corporate egress) untouched even under --yes. loopback_base_url() is the read-only probe the prompt was missing, so offboard asks only where something is really there and names the value it found. doctor diagnoses the winning entry and reports the ones it shadows as shadowed, because fixing a loser changes nothing and sends you in circles.

claude-opus-5 was not in the pricing catalog. The catalog knew claude-fable-5 and the Opus 4.x line but not the model actually serving traffic, so savings on it resolved to no price and rendered as $0 — under-reporting real work without ever failing. Now $5.00 / $25.00 per MTok. claude-mythos-5 is deliberately absent: same $10/$50 as fable, but Project Glasswing only, so it would be a guess at a model most callers cannot reach.

tests/conftest.py sandboxes the new reach. --undo and offboard can now rewrite Claude Code settings machine-wide, which is precisely why no test may keep it: a suite run from this repo would otherwise rewrite the developer's own ~/.claude — the accident that started all of this.

---

rc2 — the same failure, one layer up. rc1 fixed --always-on writing a base URL it had not proved. Three outages on one machine in a single day showed the other half: a base URL in ~/.claude/settings.json outranks the environment distil wrap sets. So when the always-on service was not running, wrap dutifully started a proxy, exported its port, and Claude Code ignored it and dialled the dead one. Every session failed with API Error: Unable to connect (ConnectionRefused) — an error naming the provider. distil wrap now refuses to start into that configuration and names the fix, and the end-of-session warning no longer guesses "agent update?" when it can check.

Distil is embeddable. from distil import compress_messages, expand_handle. The package previously exported only __version__, so nothing could depend on it without shelling out to the CLI. Not named compress/expand: distil.compress and distil.expand are modules, and Python binds a submodule onto its parent on import, so those names would resolve to the function or the module depending on unrelated import order. A test enumerates submodules and fails on any future collision.

▼413 · 0% smaller is gone. A real token count beside a rounded zero reads as broken, and it is the likeliest first-run uninstall trigger for traffic that genuinely compresses to very little. Sub-1% now renders <1% smaller with the reason, across all four surfaces. doctor also stopped reporting a routing proxy as bypassed: it inferred traffic from the savings ledger, which skips zero-saving windows by design.

A real TypeScript library. The npm package was a CLI bridge. compress(messages) now runs natively in Node and is byte-identical to the Python engine, enforced by a 92-case conformance suite. The digest tier is deliberately not ported — it mints handles into a store shared with the proxy and MCP server, and it is the tier the certificate measures. Where JS cannot reproduce Python's bytes (integral floats, integer-like keys) the port declines rather than emitting uncertified output.

Enterprise surface. OIDC + RBAC on the gateway (three ordered roles, stdlib-only JWS; RS256 refused rather than unverified when the extra is absent), a Helm chart with secure defaults asserted by test, Prometheus alert rules and a Grafana dashboard whose every expression is CI-checked against the metrics distil actually emits, distil_requests_rejected_total so quota enforcement is visible, SECURITY.md, and a security whitepaper that states the gaps — no SOC 2, no SLA, no SAML — rather than omitting them.

Reach and honesty. opencode and qwen wrap presets; Agno and Strands integrations; docs/IDE-AGENTS.md documents the proxy route for Cursor/Cline/Continue and says plainly that Copilot is not redirectable. The README now leads with what distil does and moves the proof below the fold.

Fixed while getting here: npm test ran node --test test/, which on Node 22 resolves test as a module and dies — the package's tests were not running; CI hand-listed one npm test file so new ones never executed; and cli.main() left SIGPIPE at SIG_DFL process-wide, so any later broken pipe killed pytest with exit 141 and no failing test to point at.

1.41.1 — the diagrams, and the one nobody could turn off

Documentation and accessibility only. No compressor, proxy or CLI behaviour changes.

Six diagrams were stills. Each now carries the motion its own argument needs rather than a uniform fade: ast-delta's three "same AST → unchanged" verdicts land one at a time, because the claim is cumulative — a reader has to watch reformatting fail to count, three times, before "1 / 3 defs" means anything. io sends a packet down each leg, because the point is that both directions are billed and a reader who only sees the outbound arrow half-remembers that. vision's duplicate tiles dim as they are elided, so it reads as "these two, and only these two". observability reveals its four scopes left to right, because the claim is that each is wider than the last. install is a menu rather than an argument, so its cards simply deal in. banner is deliberately restrained — it is the README hero, and the only thing that moves is the mark's three bars, which already encode compression by narrowing.

logo, logo-lockup and og stay static. A favicon that moves is a bug, and every social platform renders an OpenGraph card as a still — animating it risks the frame a scraper happens to capture being a half-drawn one.

hero-terminal.svg had twelve animations and no prefers-reduced-motion guard. It had been on the landing page since the redesign, through an entire accessibility series, because nothing checked. It is also the awkward case: it plays once and freezes rather than looping, so the usual animate { display: none } would have left it permanently blank — every group starts at opacity 0, the typed lines are revealed by zero-width clip rects, and the result box is drawn by a dash offset. The guard asserts the end state of each. A hero that renders as an empty terminal for anyone with reduced motion enabled is worse than one that animates.

Four diagrams had no accessible name at all — cache-aware, cross-sdk, domains, head-to-head. <img alt> names them on a page, but an SVG is also a document people open directly (GitHub renders docs/assets/*.svg as a page) and there alt does not exist. Titles and descriptions now travel with the files.

tests/test_svg_assets.py makes all of it permanent: well-formed XML, a reduced-motion guard on anything that moves, keyTimes that start at 0 and end at 1 and are non-decreasing, values/keyTimes counts that agree, and an accessible name. Both failure modes were verified to fail the test rather than assumed to. Static and unnamed assets are allowlisted with their reason, so the next one has to be argued for.

1.41.0 — what the cache bought, and what a name is worth

Three things of the same shape: a claim distil could make but not show, and a number that was true for the wrong reason.

distil cache — the observability half of cache alignment

distil already prevented prefix drift by construction. strategies.distil compresses the volatile tail and leaves every stable block byte-identical, so compression cannot be the cause of a cache miss. What was missing was any way to show it — and any way to name the case distil cannot fix: a caller whose own system prompt carries a timestamp.

The report deliberately mixes two kinds of number. Cache reads and writes come from the provider's own usage — ground truth about money. Prefix drift is distil's diagnosis of why, from a content-free hash of the stable blocks it sent. It never prints one without the other, because a diagnosis with no measurement behind it is a guess.

requests        3
cache reads     31,614 tokens  (billed at a discount)
cache writes    15,819 tokens  (billed at a surcharge)
hit ratio       66.6% of cacheable tokens were reads
prefix drift    1 of 2 turns changed the stable prefix (50%)

Verified live on both the streaming and non-streaming paths: the prefix held across a turn append, then a session id prepended to the system prompt forced a re-create. The turn the hash flagged is the turn the provider re-billed — derived independently, agreeing. Only system, tools and tool_choice are hashed: a conversation growing is not drift, and a warning that fires on every healthy turn is one people switch off.

The proxy now records usage_cache_read and usage_cache_create separately. Summed, they cannot tell a working cache from a thrashing one — a write is a surcharge, a read a discount, and a prefix that drifts every turn writes forever and never reads.

With no proxied requests, distil cache exits non-zero rather than printing a reassuring zero. A cache report over no data reads as a clean bill of health for a session nobody measured.

New: docs/CACHE.md, including the one cache feature deliberately not shipped — a response cache — and the three reasons why.

BFCL was being scored as prose

Its golds are bare code names; the generic matcher anchors on non-word boundaries. On the n=25 sample that credited 11 of 85 golds by accident — 'a' matching inside "tool-schemas".

Names are now matched as identifiers: a quoted JSON token, tolerating one level of escaping. The escaping is the part that makes it work — a tool schema is JSON nested inside a JSON payload, so its quotes arrive as \"base\". An earlier attempt requiring a bare "base" read 2.9% on a payload where every name was plainly present; that number measured the matcher, not the compressor. This removes the "upper bound" caveat that shipped with 1.40.0.

Support recall is now 100.0% — and the report says what that means. Visible recall is 0%: the schema sits behind a restore handle, one distil_expand away. At matched savings (89.3% vs 90.1%) truncation keeps 0 of 70 names; distil keeps all 70, none of them on screen. The table prints visible → true support: bfcl 0%→100% rather than the flattering figure alone, because a reader who assumes the model can see a schema it must first expand has been misled by numbers that are each individually correct.

Fifteen golds BFCL genuinely names a, b, c are excluded and counted. A one-letter token occurs in almost any English text, so it can be neither credited nor failed honestly — the same treatment as an abstractive answer. A recall computed against a quietly reduced denominator is an inflated result.

make gate was weaker than CI

CI ran distil validate; the Makefile did not. A local green gate promised a push it could not deliver — the one thing that target exists to prevent. Added, along with a test that compares the two by invoked subcommand and fails naming the missing one.

Also

  • Three tests fetched BFCL live to assert on --min-recall gate logic. Seven CI runners hitting the same endpoint produced an HTTP 429, and the assertion degraded into comparing its expected string against a rate-limit message. A red gate meaning "a third party throttled us" trains people to re-run until green. They run offline now.
  • eval-stack.svg still said five gates and omitted distil suite. Rebuilt animated, six gates lighting in run order.
  • phantom-file.svg and cache-delta.svg animated — the deletion is struck through and dimmed, and the two re-read bars race at true relative scale. Both resolve to their end state under prefers-reduced-motion.
  • cli.html never got a suite section from 1.40.0. Added, with cache.
  • evals.html was marking Corpus as the active nav page.

1.40.1 — the numbers, corrected

Bug fixes and honest reporting. No compressor, proxy or CLI behaviour changes; the 1.40.0 eval suite is unchanged.

distil stats crashed on a legacy Windows console. The savings line contains →, and Python's default errors="strict" turns that into a UnicodeEncodeError mid-render on a cp1252 terminal: exit 1, a traceback, output truncated. It only fired once a ledger had baseline tokens, so it had been invisible since the line was written. Streams now degrade with errors="replace" instead of failing — a reporting tool must never fail on the report.

A lifetime figure was being read as the current rate. Validating the published adoption numbers against this repo's own ledger found both correct and both misleading the same way: −20.4% lifetime against −0.4% over the last 7 days, two orders of magnitude apart, with only the lifetime number shown. distil stats now prints the recent window whenever it disagrees, names the cause, and names the remedy. A window that comes out LARGER than baseline reads +10.0% LARGER rather than the previous −-10.0%, and gets advice for overhead rather than for lossless-only.

The subscription default never said what it costs. A subscription session runs lossless-only so no digest is left unrecoverable — correct, and Tier-0 only, which measures ~0-2% against the 30-60% the recoverable digest reaches. The notice now leads with that cost and offers the persistent opt-in (distil default --modeexpand) rather than only the per-run flag. The default itself is unchanged.

Offline artifact parsing. A tool call's own ARGUMENTS no longer condemn it — a successful Write(file_path="a.py", content="No such file or directory") was recording nothing. The call boundary is parsed (paren depth, quote-aware) rather than guessed from a -> delimiter, and only a closed set of recognised invocation heads gets that immunity, so an exception repr or a result wrapper is still read as evidence. Ops inside one call share its outcome, so a failed Bash(command="rm a.py && touch b.py") records neither. The live path was never affected: it knows the real call/result boundary.

Also: --max-silent diagnostics name every component of their own total, the eval record's subject in the docs matches what the CLI emits, and an output surface that changed blocks but endangered no facts says so.

1.40.0 — what recall cannot see

Four state probes, both priced surfaces, and every result as a reproducible record.

distil retention answers "is the fact still there". distil fidelity answers the questions that survive a yes:

  • artifact state — is the file's final STATE right, or only its name present? stale (present, wrong) is reported apart from lost (absent), because a compressor that drops a whole file history is safer than one that preserves half of it: the first leaves a gap the agent can see, the second leaves a false belief it acts on.
  • overclaim — did the value keep its uncertainty? "approximately 4200 ms" → "4200 ms" is a distortion every recall metric scores perfect. Reversing a bound (at least 3 → at most 3) is reported separately as inverted.
  • continuation — does the agent still know what is left to do?
  • propagation — does a loss at turn k show up as a changed decision at k+n?

Both halves of the bill are graded: what the model READS, and what its own past answers cost when they re-enter as history.

--json now emits an eval record rather than bare metrics — schema version, dataset fingerprint, subject identity, grader provenance, and gates carrying threshold, observation and outcome. A number nobody can reproduce or compare is a number nobody should act on. The schema ships at schemas/eval-record.schema.json and is validated against real output.

Measured on the shipped reversible tier: artifact state 100% (7/7), hedge fidelity 94.7% (162/171) with 9 genuine overclaims, pending-work recall 100%, output surface 6 blocks digested with 4 facts removed and 0 silent failures. CI gates at --max-silent 15 — the measured band, not zero, because gating at zero would assert a property the compressor does not have.

The rule these probes enforce is that no gate may pass without evidence, and the cross-audit found it broken in six places — each one green, each one measuring nothing: an empty gate list, a fresh shadow ledger, an unscored output surface, a propagation profile with no events, a threshold accepted without running, and a probe with a zero denominator. All six now fail rather than certify. Three sweeps close the classes underneath: every alternative of every alternation must fire, no gate may certify an evidence-free input, and every distil … command printed in the docs must actually parse.

Also in this release: the SDLC auto-fix loop closes end-to-end (#74–#80). A GITHUB_TOKEN push raises no workflow events, which had silently broken three chains.

Upgrading is safe and additive — no behaviour of the compressor, proxy or CLI changes. distil fidelity is new; every existing command keeps its output shape.

1.39.1 — the release path, hardened and then actually run

The installed package is unchanged. Every commit since 1.39.0 touches .github/workflows/ or packaging/homebrew/distil.rb, neither of which ships in the wheel — distil/ and the bundled corpus are byte-identical to 1.39.0. Nothing about compression, the proxy, or the CLI is different, and there is no reason to upgrade for behaviour. It is cut so the hardened publish path executes against a version that is not yet on PyPI, npm, or the tap, rather than waiting to find out during a release that matters.

Fixed — publish jobs that reported success while doing nothing

  • A missing NPM_TOKEN now fails the release instead of exiting 0 behind a warning. That silent skip is how npm sat on 1.25.0 while PyPI reached 1.39.0 — fourteen releases, each green, publishing nothing, while the README carried an npm badge and an npm i distil-llm instruction aimed at the stale build. The registry now holds exactly 1.25.0 and 1.39.0, which is the evidence.
  • A missing TAP_TOKEN now fails the release for the same reason, behind an even quieter ::notice::. It worked only because the token happened to be set.
  • Both check idempotency before demanding a credential, so re-cutting a release with nothing to publish neither fails nor needs one. For the tap that comparison covers the url, the sha256 and the version: a re-cut tag keeps the version and changes the hash, and matching on version alone would have skipped repairing a formula whose brew install fails its checksum.
  • A tarball fetch failure now says "does the tag exist?" instead of a bare exit 56 — the guard existed but was unreachable under set -euo pipefail.

Fixed — auditors that never saw the fix

  • Both cross-audits now run on synchronize. They triggered on open only, so an auditor reviewed the commit that contained a bug and never the correction, unless someone applied the agent:audit label by hand. Bot pushes are skipped only on synchronize (the fix loop's rounds); PRs opened by Renovate, Copilot and release bots still audit, because a blanket bot exclusion silently withdraws coverage from dependency and generated-code changes.
  • Concurrency is job-level, not workflow-level. A workflow-level group is joined by the run before the job's if is evaluated, so an event the job will skip still takes a slot — cancelling the audit in flight, and evicting a queued one, since GitHub keeps a single pending run per group. The PR ends with no verdict, which reads exactly like a clean one. A skipped job never joins the group at all.

1.39.0 — certify the HTML transform, and a headline one fixture cannot swing

Documentation

  • langchain-distil is now named in this repo. LangChain merged langchain-ai/docs#4943, listing the package in its community middleware integrations with docs_url pointing here — and the README, docs/langchain.html and docs/integrations.html mentioned it exactly zero times. A reader clicking through from LangChain landed on a page that never named the package they had just been told to install. All three now carry install and usage for the real entry points (compress_messages, pre_model_hook, as_runnable), and three contract tests pin the symbols the published wrapper imports, so renaming one can no longer break it silently.
  • The corpus is eight domains, and the site said seven — 42 stale claims across 16 files, caused by this release's own web-research trajectory. domains.svg gained the row (89.8%, PASS, 428 pruned), the viewBox to fit it, and the animation stagger the other seven have.

Added

  • An HTML trajectory in the certification corpus (corpus/web-research.json). 1.38.0 shipped the HTML transform default-on in the serving path while every corpus trajectory carried logs, JSON or prose — so distil bench never routed a byte through compress/htmlx.py and the transform was certified by nothing. ADR 0003/0004 set the precedent that a new content type is certified before it goes default-on; this pays that debt.

A 4-turn web-research agent whose tool results are real HTML documents: chrome the extractor must strip (script, style, nav, cookie banner, aside, footer) wrapped around an <article> carrying the DECISION the oracle grades. It certifies at 89.8% savings, and it has teeth — an extractor patched to swallow <article> drops match_rate to 0.0 and fails the gate.

Changed

  • The retention headline is now a per-domain macro average, so one fixture cannot swing it. distil retention reports 21.4% — the mean across all 8 domains, each counted once — and prints the fact-weighted figure beside it rather than as the headline.

The reason is the previous entry. Adding a single HTML trajectory moved the fact-weighted number from 9.8% to 62.6% with nothing about the compressor changing: web-research alone contributes 666 of the 1083 facts at 4.4% visible, because HTML is dense in href URLs that extraction drops as navigation. A number that a fixture can move that far is measuring the corpus, not the compressor.

Both are still reported, because the two diverging is itself a signal that the corpus is unbalanced — and the report gained per-domain rows, which is the read that actually shows the spread (0.0% on coding, 95.6% on web-research). The JSON payload leads with macro and labels micro with why it moves. --max-lost is unaffected: a lost fact is a count, not a ratio.

1.38.0 — recall, and a number a stranger can check

Every quality gate distil shipped until now was graded on our corpus against our oracle. The statistics were rigorous; the external validity was zero — a reader had nothing they already trusted to check us against. And none of it answered the question users actually ask, which is not "did the next action survive?" but "is the number I needed still there?"

Added

  • distil retention — fact-level recall. Keyed numerics, artifacts (paths/URLs/hashes/UUIDs) and error lines are classified retained / recoverable / lost. recoverable is verified, not assumed: the handle must appear in that block's own compressed text and the fact must appear in that handle's restore bytes. On the corpus: 100% true recall, 90.2% visible, 0 lost — so being reversible rather than lossy is worth 9.8% recall, the first time that property has had a number instead of an argument. Now a per-commit CI gate (--max-lost 0); it is zero-cost, needs no API key and no network.
  • distil retention --dataset {hotpotqa,squad} — public ground truth. Graded against answer keys written by someone else, next to a truncation baseline tuned to reproduce distil's own token savings on the same case. HotpotQA (n=100, 14.3% savings): 100% answer recall and 100% gold-sentence recall, vs 91.6% / 82.7% for truncation at 14.1% savings. Rows load from the HuggingFace datasets-server REST API as plain JSON, so the core stays stdlib-only (dependencies = [] holds) and cache under $DISTIL_HOME/datasets for offline reproduction. Runs nightly, not per-commit: a required check gated on a third-party host would make main's health hostage to their uptime.
  • distil retention --live — recall on your own traffic, content-free. A sampled meter (--retention RATE; 0.05 by default under wrap, 0 under a bare proxy) scores retention in-process and persists counts only — three integers per dimension, with deliberately no field that could carry content. There is no plaintext session recorder to secure, which is the whole point. Bounded to 64 KiB per request, fail-open, and unlike --shadow it costs no extra tokens because there is no second upstream call.
  • Reversible HTML extraction — the web-fetch blind spot. Building the retention harness exposed a capability gap it could then measure: distil compressed 0.0% of an HTML tool result. The cause was structural, not a missing heuristic — minified markup arrives as one enormous line, so Tier-1's line-folding had nothing to fold and the JSON/record folds do not recognise markup. Any agent with a fetch or browser tool was paying full price for <script>, <style>, nav and footer chrome.

distil/compress/htmlx.py extracts the content with stdlib html.parser and keeps the exact original behind an expand handle. Measured on real pages:

pagebeforeaftersavedfacts lost
Wikipedia article281,093 tok14,260 tok94.9%0
Python docs page32,322 tok4,229 tok86.9%0

Deliberately recall-biased: only tags that cannot hold article content are dropped, plus four unambiguous chrome landmarks (nav/footer/aside/form). <header> is kept because it usually wraps the <h1>, and img alt text is kept because it is often a figure's only description. Unlike a lossy extractor the heuristic is recoverable — distil retention measures 100% true recall with 0 lost on the pages above, so a bad call costs one distil_expand, not the content. Skipped in verbatim mode (no expand tool to recover with) and on recency-exempt blocks (the agent's latest output stays byte-exact).

Documented tradeoff: on raw HTML most probe-able "artifacts" are href URLs, which extraction drops as navigation. They stay recoverable, but visible recall for that dimension is low by design — now a measured number instead of an unknown.

Fixed after cross-audit

All three findings from the PR's independent audit reproduced, so all three are fixed:

  • Unclosed chrome tags swallowed the article. Chrome skipping keyed off a matching close tag and real HTML often never sends one, so content before an unclosed <aside> was emitted while the article after it was dropped. An <article>/<main> landmark now ends any active skip, plus a skipped-data budget as a backstop for <div>-built chrome. script/style stay exempt — their payload is legitimately huge.
  • Synchronous disk I/O in the request path. RestoreStore.expand() falls back to a disk read for unknown handles, and the meter called it for every handle= match — so tool output merely containing that string could turn one sampled request into a series of file reads. Matched handles are now intersected with the store's in-memory set.
  • Short answers false-matched in recovered bytes. "12" matched inside file-12.csv. Recoverability now requires token boundaries, and values under four characters require whitespace boundaries.

That last fix then failed the corpus gate, which is how it proved its worth: it exposed that _NUMERIC_RE had been extracting values ending mid-token — invoices=88 clipped from invoices=88ms, junk like T09:14 from inside a timestamp. Extraction now takes the whole token including its unit. The probe set is smaller and cleaner (417 facts, not 503), which moves the corpus figures to 90.2% visible / 100% true / 0 lost and the measured value of reversibility from 12.3% to 9.8%. The earlier number was inflated by junk probes; this one is honest.

Honesty notes

Three defects found while building this, each of which had been reporting numbers in distil's favour:

  • HotpotQA's comparison questions answer "yes"/"no", which is never a span in the passage. Grading them reported the dataset's answer format as 10% compression loss. Answers are now graded only when present in the uncompressed context.
  • SQuAD v2's unanswerable questions (55% of the split) were being scored as retained — a free win on a majority of cases. They are excluded and counted separately.
  • A recall number from compression that barely engaged is arithmetic, not evidence, so --min-recall now fails below 1% savings. This is not hypothetical: distil's compressors correctly decline to touch short prose, so --shape prose yields 0% savings and a vacuous 100% recall. The default --shape json reflects how a retrieval tool actually returns documents.

Worth knowing: on SQuAD at 84.8% savings, true recall is 100% but visible recall is 0% — every gold answer sits behind a handle. Lossless, but one round trip per answer. The report says so instead of printing only the reassuring number.

1.37.0 — you will be asked, once, at the moment it makes sense

The census undercounted because nobody was asked. distil onboard was the only place that ever requested consent, so anyone who ran pipx installdistil-llm and went straight to distil wrap — the natural path, and the one the README showed most prominently until this week — was never asked and could never be counted. The live numbers made the shape of it obvious: 2 opted-in installs against 11,102 monthly downloads and 374 unique cloners. That is not a counter that is broken; it is a question that was never put.

You will now be asked once, at the end of a session that actually saved something. That is the only moment the question is fair: you have run the thing, there is a real number on screen, and "share this?" is concrete rather than hypothetical. The prompt sits before the census send, so a first yes counts the session you were looking at rather than one a day later.

Every guard the original prompt established is kept, and one is new. Asked once — a stored yes or no ends it permanently. DO_NOT_TRACK and DISTIL_NO_TELEMETRY win outright. Both streams must be a terminal, so pipes, CI and headless runs never prompt and therefore never enrol. Ctrl-C is not an answer — recording a decline there would burn the one chance to ask on a keystroke that meant "not now" — though unanswered prompts stop after three, because a prompt nobody answers must not become a prompt nobody escapes. The new one: nothing is asked if the session saved nothing. Asking someone who got no value to share their numbers is a worse question and a worse experience than staying quiet. And none of it can raise: this runs on the wrap teardown path, where an exception would change your exit code over telemetry.

What this does not do is make you identifiable. install_id is a random UUID, and distil census off deletes it — so a machine that toggles consent mints a new one. There is no way to tell one person's second laptop from two different people, and there will not be: the fix for a small opted-in set is a larger opted-in set, not fingerprinting.

Docs. A light theme (page chrome only — diagrams and code blocks stay dark, because SVGs loaded through <img> cannot inherit page CSS and half-recolouring them would look broken rather than deliberate; contrast measured at 7.24:1 body and 17.5:1 headings). Per-page Copy as Markdown, because these docs are read by agents as often as by people. And eight per-framework pages — Anthropic SDK, OpenAI SDK, LiteLLM, LangChain, Vercel AI SDK, Agno, Strands, CrewAI — each carrying the caveat that actually bites, such as CrewAI resolving planning_llm separately and silently bypassing the proxy.

1.36.0 — the image transform is certified, and now on by default

1.35.0 shipped vision compression merged but inert: the gate demanded a certificate and none could exist, because the certification path was text-only end to end. Block.text is a str, so a trajectory could not hold an image, and every runner graded a decision from text. This closes that.

Certified, live. Block now carries optional media; the Anthropic runner renders it as real provider image blocks; vision is a registered strategy; and corpus/vision-ci-dashboard.json is a decision-bearing vision trajectory with real, CRC-valid PNGs whose context accumulates, so byte-identical screenshots genuinely pile up. Against claude-opus-4-8: 100% decision-equivalence, A/A self-agreement floor 100%, TOST p<0.0001, VERDICT PASS.

The first live run failed, and that is the point. Turn 3 diverged — baseline promote_release, compressed open_failing_build. The cause was not the compression: the runner was hoisting every image to the front of the turn, severing each screenshot from the caption identifying it. An A/B whose two arms differ in prompt shape measures the shape. Interleaving fixed it. No offline deterministic oracle could have surfaced that, which is exactly why this domain is certified live and is deliberately excluded from distil bench.

What changes for you on upgrade. The maintainer's certificate ships in the package, so vision de-duplication is enabled by default. An agent that sends repeated identical screenshots — UI automation, dashboard polling — will start seeing duplicates elided, and its reported savings will rise accordingly. Only byte-identical payloads are ever elided, the first occurrence is untouched, url sources are never treated as duplicates, and every elision is recoverable through distil_expand. DISTIL_VISION=0 disables it outright.

What you are inheriting, precisely. The shipped certificate states its own scope: the model, the corpus, the verdict, and the exact command to reproduce it. It does not certify your traffic. Run distil certify --strategy vision--runner anthropic against your own captured trajectory to certify your workload — and if that run FAILS, your result wins: a local failing verdict is never overridden by the shipped pass. The gate reads three sources in order (DISTIL_VISION → local → shipped) and every one of them must parse and carry an explicit passing verdict. Inheriting a certificate is not skipping the gate.

1.35.0 — see what it does before you trust it

Eight merged PRs. The theme is the same one twice: a number that claims to be current, and a preview that claims to be complete. Both were wrong, in public.

Three adoption numbers were stale, not false. Reported as "the page shows a previous version when I have the latest" and it turned out to be three separate bugs of one shape. distil wrap is deliberately long-lived — hot-swap replaces the proxy worker on upgrade so your session never restarts, which means the wrap parent keeps running the code it started with, for days, and it is the process that emits the census. Every wrap user's beat reported whatever was installed when their session began; a machine running 1.34.0 was observed beating 1.28.0, two upgrades later. by_version also counted every install ever seen under a chart that says "in the wild", so a machine that pinged once and vanished sat there forever. And the live hero read "2 machines saving now" directly above its own sentence "summed from 1 machine that actually reported savings" — the endpoint's active field is a consent count, and the page fell back to it. All three fixed; the version now reads from disk, the histogram is 30d-scoped, and "saving now" means saving.

distil simulate — a dry run that says what it will NOT touch. Running the real pipeline locally with no model in the path was already possible. What was missing is the question that actually decides adoption: what would you leave alone, and why? distil already computes that — the recency exemption, the assistant's own words, tool_use arguments, lines the keep policy pins as decision-bearing — so the protected set is now first-class output, per block, naming the rule. Two guarantees are reported separately, because conflating them made the first version lie: byte-exact (not touched at all) and lossless-only (bytes may change, meaning preserved, nothing moved behind a handle). Recency is the second kind, not the first — it exempts a block from the Tier-1 digest, not from Tier-0.

Prometheus and OpenTelemetry. GET /distil/metrics serves the standard text exposition, stdlib-only, behind the same admin gate as /distil/stats because the series are labelled by tenant. OTel counters mirror it, recorded at the same instrumentation point as the span attributes and before the span guard, so they survive tracing being sampled off. This corrects two published claims: the README and Deploy & Security both said distil has no metrics endpoint and cited that as a differentiator. It has one; the gate is the differentiator, and it is tested.

Images become a certifiable content type (ADR 0003). The prevailing technique here downscales, which is lossy by construction — the model sees a different image and nobody can say whether the answer changed. Instead, the second and later appearances of a byte-identical image become a recoverable reference; the first is untouched. It ships disabled: vision.enabled() parses the certificate and requires an explicit passing verdict, so with none present the adapter is byte-for-byte what it was before. Reversibility is proven; decision-equivalence is not yet, and ADR 0004 records exactly what blocks it.

A savings figure that was wrong in public. The proxy's estimator never counted image blocks, so an image was ~0 tokens on the "before" side. Any agent that sends images had an understated baseline. Now counted at pixel-area cost — which means reported savings and census figures will shift for image traffic. That is a correction, not an improvement.

Docs search. 22 pages with no way to search them. ⌘K, static index, no service and no dependency. A test fails if the committed index goes stale, because a generated artifact that is committed rots the first time someone adds a page — and rots silently, since search simply never returns it.

Also: Agno, Strands, Cursor and CrewAI added to the integration matrix, each with the caveat that actually bites (Cursor's override covers its agent panel but not tab-completion; CrewAI resolves planning_llm separately, so leaving it unset bypasses the proxy with no error). ADR 0004 stack-ranks what is genuinely left and records three gaps the earlier audit had understated — we had more than we thought. And the Gemini cross-audit workflow, dead on every PR, now runs from the base commit with a fail-closed tool allowlist rather than trusting the PR's tree.

1.34.0 — say whose savings those are

The counter fix in 1.33.1 stopped the community total moving backwards. This is the other half of the same problem: what that total actually means.

Consenting is not contributing. The adoption page counted every machine that had turned the census on as a machine that was saving. On the real data that made "2 machines saving now" out of one machine with 1.6B tokens and one that had reported tokens_saved: 0 twice and never run anything. The rollup now emits contributing (installs that have actually reported savings) alongside instances, and the page uses it: the headline reads "from 1 machine", the install tile reads "1 of 2 have reported savings · 1 opted in but idle", and below five contributors the hero carries an explicit small-sample notice — this is one real ledger, not a community aggregate. Aggregates written before the field existed make no claim rather than falling back to the consent count, which is the overstatement this removes.

Your savings, one number. distil dashboard --web computed savings as lifetime-raw × current calibration factor, the exact method census.py documents as wrong: it restates every token already earned whenever calibration refines, so the number on screen can drop without a token being un-saved. It also disagreed with the census on the same machine — 1,522,590,876 against 1,491,058,879. Both now read the count-time accrued total through census.accrued_tokens(), so your dashboard, your census payload, and the community rollup are the same figure.

Most users are never asked. Consent is offered in exactly one place — inside distil onboard, at a TTY. Anyone who went from uvx/pipx straight to distil wrap was never asked at all, which is most of them. distil stats now mentions the census once, only to someone with real savings to contribute, only at a terminal, never twice, and silent under DO_NOT_TRACK. It does not send anything and does not grant consent — ignoring it leaves you un-asked.

1.33.1 — the zipapp can name itself

distil.pyz — a release asset offered as the install path for anyone PyPI is blocked for — reported its version as 0+source rather than the version it was built from. It has done so since at least 1.31.0.

The archive carries no dist-info, so importlib.metadata misses; and the pyproject fallback in distil/__init__.py calls read_text() on a path inside the zip, which a zipapp cannot open. Both paths failed silently to a literal. build_pyz.sh now stamps distil/_version.py from pyproject.toml at build time — zipimport can import a module even though it cannot read a file — and the fallback chain tries the stamp before the unreachable pyproject read.

Nothing functional changed: the archive always executed correctly, it just could not answer --version. Customer-facing regardless, since that string is what a bug report quotes.

Why it survived three releases, which is the more useful finding: no test ever ran the .pyz. And running it naively still would not have caught this — under any interpreter with distil-llm installed, importlib.metadata answers correctly and the zipapp path never executes, so the artifact looks green in exactly the environment a developer tests it in. tests/test_packaging_smoke.py now builds the archive and runs it in a bare venv --without-pip, where the package is genuinely absent. The test was verified to fail (zipapp reports'distil 0+source', pyproject says '1.33.0') with the stamping step removed.

Also folds in the in-repo Homebrew formula sync, so the tag and main no longer diverge — v1.33.0 was tagged one commit before it.

Census: the community total could run backwards, and Ctrl-C burned the consent ask

saved is monotonic per step, but it was advanced from two places (the daily census and the near-real-time heartbeat) across several concurrent processes — wrap, proxy worker, gateway, webdash — each doing load → step → write-whole-file with no lock. The last writer clobbered the rest: a process holding minutes-old state wrote a smaller saved back and rewound raw_seen with it, so an already-banked delta could be counted twice. That is how a total which can only rise published 1.44B, then 1.33B, then 1.16B, then 1.48B. _savings_locked() now holds an exclusive flock across the whole read-modify-write, so every writer steps from the newest state.

Separately, except (EOFError, KeyboardInterrupt) around the consent prompt fell through to opt_out(). A user who pressed Ctrl-C — or whose stdin was a pipe — was permanently recorded as having declined and was never asked again. No answer now leaves consent unset, so the question can be asked another time.

1.33.0 — the reports are readable now

A WCAG 2.2 pass over the docs site and every HTML surface distil generates — the gateway dashboard, the live web dashboard, the savings ledger, the technique leaderboard, the benchmark report, and the dissect portal. Contributed as eight focused PRs by @pjdoland (#34–#41).

Some of this is plain bug-fixing that happened to surface through an accessibility lens. Two docs pages wired their mobile navigation button to a function that was defined nowhere, so the menu threw a ReferenceError for everyone; nine code blocks had a copy button pointing at a missing copyCode. (The report that the working copy control captured the word "copy" rather than the snippet did not reproduce under an intercepted clipboard.writeText — the nine dead buttons were real and exactly counted; that secondary claim was not.) The leaderboard's "not certified" state was an em-dash at 1.75:1 contrast — a verification marker you effectively could not read, in a project whose entire pitch is that you should verify rather than trust. Muted text across the reports sat at 3.28:1; it is now 5.45–6.37:1, comfortably past AA.

The interaction changes are real UX wins, not just conformance. The gateway dashboard and the sessions portal used <meta http-equiv="refresh">, so they reloaded wholesale every 5 s and 15 s — destroying scroll position, text selection, and keyboard focus while you were reading them. Both now poll JSON and patch rows in place, with a visible Pause control (WCAG 2.2.1/2.2.4). Charts in the dissect report carry accessible names and a data table fallback with the same numbers. Tooltips are focusable, dismissible with Escape, and announced. Tables have real headers, captions, and scope; the docs sidebar is a labeled <nav> with list semantics and aria-current.

No compression behavior changed: make gate (corpus non-inferiority + byte fidelity) and the full suite pass unchanged.

Provider compaction, experiment 2 — the default clearing policy

1.32.0 measured Anthropic context editing at keep=0 and found 92.5% of agent decisions changed. The obvious rebuttal is that nobody runs keep=0. So this release runs the shipped default, keep=3, on a 7-round corpus whose two decision-bearing tool results sit at the head and whose four routine rounds (inventory, shipping scans, comms, promotions) sit at the tail — the ordinary shape of a long agent trajectory, where the load-bearing facts are gathered early and then buried.

Retention did not help. Clearing changed 95.0% and 100% of decisions across two independent executions of the pre-registered protocol (both published; selecting one after the fact is the practice this harness exists to refuse). OpenAI's summarizing compaction changed 20.0% on the same corpus — a ~5× gap, consistent with the ~7× at keep=0.

The rate is not the finding. The failure mode inverted: at keep=0, 37 of 40 flips were tool→text — the agent lost its facts and stopped acting. At the default keep=3, stalls nearly vanish and 29–31 of 40 flips become tool→wrong tool. The three surviving routine results are enough to keep the agent confidently acting while the records it needed are gone. Keeping the most recent tool uses cannot protect a decision that depends on the most relevant ones; it converts a visible stall into a silent wrong write. Neither feature certifies decision-safe at α=0.1.

certify-provider runs are now resumable. A 360-call run takes ~40 minutes and OpenAI's compaction path returns 500s in bursts; two full runs were lost to it before each finished case was made durable. Every completed case is appended to cases.jsonl stamped with the protocol hash, so a re-invocation replays what it already bought and a changed parameter orphans the ledger instead of blending two experiments. A resume reuses the original protocol.json, so "pre-registered before the first call" stays literally true across the restart, and calls_made reports the whole experiment's cost rather than the last attempt's.

1.32.0 — certify the provider's own context manipulation

distil certify-provider: a pre-registered, A/A-controlled, budget-capped live A/B that measures whether the provider's own context manipulation changes an agent's next decision — Anthropic context editing (clear_tool_uses) and OpenAI server-side compaction (--provider openai). Vendors will not publish decision-equivalence for their own features; a third party on the wire can.

The design is the shadow/certificate machinery pointed at a new A/B: same multi-turn tool transcript, manipulation ON vs OFF, plus a second baseline arm for the sampling noise floor. Firing is ground-truthed per request (applied_edits / the compaction output item); cases where the manipulation never fired are excluded from the sample. The protocol (n, α, δ, votes, trigger, model) is written to disk before the first API call, live calls are hard-capped, and transient upstream failures retry with bounded backoff instead of burning an unattended run.

First pre-registered certificates (n=40, majority-of-3, worst-case configs on a synthetic decision-bearing corpus, benchmarks/results/provider-compaction/): Anthropic clearing changed the agent's decision on 92.5% of fired cases — 37/40 flips were tool→text, the agent stops acting rather than acting differently — while OpenAI compaction changed 12.5%. Neither certifies decision-safe at α=0.1. Honest scope: aggressive triggers on transcripts built so tool results are decision-bearing; the default-config number on long real traffic is future work.

Also: the live extra now includes the OpenAI SDK.

1.31.1 — prove the attestation gate

No runtime change. This release exists to run the corrected attestation check in CI, because a verification step that has never executed is not a verification step.

1.31.0's release job failed on a false negative: the gate read /pypi/<pkg>/<ver>/json, which PyPI does not populate with attestation data. Every release back to 1.19.0 does carry a bundle, on the /integrity/<pkg>/<ver>/<file>/provenance endpoint. The fix landed in 95fd3f5 but not in the v1.31.0 tag, so re-running that job replayed the same bug — a rerun could never have gone green.

This tag carries the corrected gate. If the release job passes, the check works against a live publish; if it fails, the check is still wrong and we find out now rather than on a release that matters.

1.31.0 — evidence that checks itself

Four things that all failed the same way: a claim nothing verified.

PEP 740 attestations — now enforced by the release job

README.md claimed releases carried Sigstore attestations. They did — every release back to 1.19.0 carries an attestation bundle. What was missing was any check, so the claim was merely unverified rather than false.

The release job now queries PyPI for the version it just published and fails if no attestation bundle is present, so the claim cannot drift.

Correction. An earlier draft of this entry — and commit b3eb4a7 — stated that every release published with attestations=none. That was wrong. It read /pypi/<pkg>/<ver>/json, which PyPI does not populate with an attestations field; the data lives on the /integrity/<pkg>/<ver>/<file>/provenance endpoint. The first run of the new gate failed 1.31.0 for the same reason — a false negative in the check itself. Both the gate and the claim are corrected here. The irony is not lost: a verification step that reported a false failure is the exact defect class this release is about.

Certificates name the oracle that graded them

A Certificate recorded α, δ, n, savings and a guarantee — but not what produced the losses. A run graded by the synthetic DECISION: string-match oracle was byte-identical to one graded by claude-opus-4-8. Certificate.grader is stamped from the runner, and the synthetic oracle can never read as a model: Graded by: deterministic (synthetic DECISION:oracle — NOT a model).

Per-request receipts

distil receipts — a hash-chained, content-free record of what happened to each request: counts, mode, handles issued, whether they still resolve, and the certificate that authorised the mode. Edits, deletions and reorders are all detected. Receipt.FIELDS is the exhaustive persisted set and a test fails if anything outside it reaches disk; no prompt or completion text is ever written. Emitted for every 2xx the proxy serves — not only inside a wrap session, and not only when a savings ledger happens to be attached.

Fixed: digest reported changed=True on byte-identical output

tier1.digest() returned True unconditionally. Where every line is must-keep — a test log whose verdict policy pins each PASS line — nothing is dropped, no marker is emitted, and the output is identical to the input. Callers believed it anyway: RestoreStore._record persisted ~19KB of plaintext per block to ~/.distil/restore for content that was never digested and can never need recovery (three such entries on a single request), and the MCP server handed back a handle for text it had not compressed. changed now means the output actually differs.

Found while chasing a receipt that read 27764->27764, saved=0 beside three issued handles. The savings header itself was honest — the verdict keep policy was correctly retaining every line — so no reported savings number was ever overstated.

Also

  • Packaging gate extended to Docker (ENTRYPOINT/CMD resolve to real targets, plus a CI job that builds the image and runs it), the Homebrew formula (self-consistent url/version/ sha comment), and the Claude Code plugin — which caught plugin.json sitting at 1.8.6, 22 releases stale.
  • release.sh now aborts a tag when pyproject, CITATION.cff, plugin.json and server.json disagree on version.
  • server.json description brought under the registry's 100-character cap.

1.30.0 — the official MCP registry entry actually launches

distil has been listed in the official MCP registry since 2026-07-17 with a launch spec that could never have worked:

$ uvx distil-llm mcp        # what the registry entry resolves to
An executable named `distil-llm` is not provided by package `distil-llm`.

The distribution is distil-llm but its console scripts were distil and distil-mcp, and uvx <pkg> runs the executable named after the package. Anyone who discovered distil through the registry — the highest-traffic MCP discovery surface — and used its declared spec got a failure.

uvx distil-llm now starts the MCP server

Adds a distil-llm console script pointing at distil.mcp_server:serve. Because serve() ignores argv, the already-published uvx distil-llm mcp form works too — so this repairs the live registry entry for existing clients without waiting for anything to be republished. distil, distil-mcp, and uvx --from distil-llm distil-mcp are all unaffected; every form was verified against a built wheel.

Ships server.json at the repo root so the registry entry has a committed manifest to publish from, instead of existing only as server-side state.

Fixed

  • scripts/release.sh rewrote url, sha256 and version in the Homebrew formula but not the comment naming which tag that sha belonged to — it sat at v1.11.2 while the hash moved 18 releases on. A comment that confidently names the wrong tag is worse than no comment: anyone recomputing the hash would have checked it against the v1.11.2 tarball and concluded the formula was corrupt. It now moves in the same sed.

1.29.1 — nothing can pin a worker thread forever

Two unbounded waits, at opposite ends of the same pipe. Both let one stuck peer outlive a SIGTERM'd worker; both are now finite, for the reason the upstream socket was already finite.

The drain has a deadline

server_close() joins the non-daemon handler threads — that join is how in-flight streams finish draining — but it had no bound, so a single wedged handler pinned the whole worker. Caught as an intermittent macOS CI hang, and captured in a stack rather than inferred:

handler thread : distil/proxy.py _post_upstream -> socket readinto
main thread    : socketserver.py server_close -> join -> threading join

server_close() now runs on a helper thread joined with _DRAIN_BUDGET_S (DISTIL_DRAIN_BUDGET_S, default 300s — well under the supervisor's 15-minute SIGKILL cap). Bounding the join alone is not enough: returning from main() re-joins those same non-daemon threads at interpreter shutdown, so past the budget — shadow and savings already flushed — the worker exits directly. Only the hot-swap worker could ever hit this; QuietHTTPServer inherits daemon_threads = True, so the in-thread proxy's join is a no-op.

Client sockets have a timeout

_DistilHandler inherited StreamRequestHandler.timeout = None, so accepted client sockets had no timeout at all: a peer that connects and goes silent, or stops reading while the response fills the socket buffer, parked its handler thread for the life of the process. The upstream socket has carried a finite timeout since it was written, with a comment saying it is finite precisely so a wedged upstream "can never pin a worker thread forever" — the client half now gets the same bargain via _CLIENT_TIMEOUT (DISTIL_CLIENT_TIMEOUT, default 600s, generous because HTTP/1.1 keep-alive means an idle agent between turns is sitting in exactly that read). socket.timeout joins handle_error's quiet list, since a stalled peer now surfaces there on write the way a vanished one already did.

Measured, both directions, same scenario — connect, then say nothing: timeout=None still held the connection open past 8s; timeout=2.0 closed it at 2.0s.

Fixed

  • Test teardown called upstream.shutdown() without server_close(), which stops the accept loop but leaves the listener open — a worker connecting afterwards completed its handshake into the backlog and waited out the full 600s upstream timeout for a reply nobody would send. "Upstream is gone" was a black hole rather than ECONNREFUSED. New _stop_upstream() helper, applied at all 12 sites.
  • CI now gates on ruff format --check (whole tree, pinned ruff@0.15.10). The formatter had silently drifted on 16 files because only ruff check was gated; that drift is also fixed here, verified AST-identical so nothing changed but layout.

Gates

1832 tests · ruff · format · mypy · bench · verify · validate — all green, and both fixes carry a regression test that fails without them (test_drain_is_bounded_when_a_handler_cannot_finish, test_stalled_client_cannot_pin_a_handler_thread).

1.29.0 — the MCP server tells agents what it actually does

Distil's MCP server worked but under-described itself: three tools with one-line descriptions, no annotations, and no install instructions anywhere. An agent had to infer whether distil_expand was safe to retry, and a human had to guess the client config.

MCP tools are annotated and self-describing

All three tools now carry MCP annotations — title, readOnlyHint, destructiveHint, idempotentHint, openWorldHint — so a client knows distil_expand is a safe, repeatable, offline read without parsing prose for it. Descriptions state what was previously implicit: the JSON return shape, the literal error string on an unknown handle, that the store is local and encrypted, that handles age out after DISTIL_RESTORE_TTL_DAYS, and when not to call each tool. distil_expand also declares pattern: ^[0-9a-f]{8}$, matching the handle regex the server already enforced.

Glama's tool-definition rubric scored distil_expand 3.7/5 ("no annotations provided, so description bears full burden"); since the server-level score is 60% mean + 40% minimum, the weakest tool set the ceiling.

Install instructions that fit on one screen

README gains an MCP section and docs/integrations.html is rewritten: the one-line claude mcp add distil -- distil mcp, the shared mcpServers JSON for Claude Desktop / Cursor / VS Code, a uvx no-install variant, a client-free tools/list verify command, and a per-tool "when your agent reaches for it" table. Both state the thing users get wrong — this is the recall path, not the savings path: distil wrap compresses traffic, the MCP server exists so any agent can expand a handle it finds in context.

Fixed

  • CITATION.cff was stale at 1.27.0 while pyproject.toml shipped 1.28.0. The v1.28.0 tag was cut without scripts/release.sh, whose consistency check (release.sh:94) exists precisely to catch this. Both are now 1.29.0 — cite the version you actually installed.
  • glama.json added at the repo root so the registry listing stays claimed if the repo ever moves to an organization.

1.28.0 — recoverable compression everywhere, and it streams

Closes the "expand-injection gap." Lossy Tier-1 digest could leave stubs the agent couldn't recover, and the recoverable path lost streaming. Both are fixed: recoverable digest is now the default on every path and keeps time-to-first-token.

Streaming distil_expand interception — recover WITHOUT losing TTFT (new)

Recoverable digest injects a distil_expand tool so the model can pull an elided block back on demand. Handling that used to require buffering the whole response (the expand loop needs the complete turn), turning time-to-first-token into time-to-last-token on any session that carried a digested stub. distil/streamexpand.py now speculatively streams: it relays tokens as they arrive and only intervenes if a distil_expand call actually appears — suppressing that internal call, resolving the handle, re-querying, and splicing the continuation into the same client stream (re-indexed, one coherent message). Most turns never call expand and stream untouched; the rare expanding turn still streams its answer. The agent never sees distil's recovery tool.

Recoverable by default on every metered path

make_app now turns the expand loop ON wherever lossy digest will actually run (any metered/PAYG session that didn't force --verbatim). Previously a bare distilproxy/wrap without --expand, the async proxy, or a direct make_app caller could run Tier-1 digest with no recovery tool — leaving irreversibly-lossy stubs, the exact harm the subscription force-verbatim already prevents, silently un-guarded on PAYG. With streamexpand there is no TTFT reason to leave that gap open. --verbatim still wins; subscription stays lossless-only (no digest).

The async proxy stops emitting unrecoverable stubs

distil.aproxy injects no expand tool and runs no expand loop, so any Tier-1 stub it created could never be pulled back. It now folds all lossy digest into verbatim — the async path stays Tier-0 (reversible lossless transforms like the #24 columnar fold still apply, so real savings remain) and never leaves an irrecoverable stub. (Recoverable Tier-1 on the async path awaits a streaming expand loop of its own.)

Gates

  • Full suite green (1828 tests; +8 for streamexpand — pass-through, interception+splice, block re-indexing, SSE frames split across read boundaries, usage summing, max-iters bound, re-query failure — plus an end-to-end streamed intercept through the real proxy). Coverage ≥95%; pinned ruff + mypy clean; distil verify / bench / validate unaffected.

1.27.1 — the shadow gate actually runs; the compression mode stops flipping

Two bug fixes in the request path's measurement and policy layers. Neither changes what the model is sent on a normal request — but both were silently degrading guarantees distil advertises.

The decision-equivalence shadow gate was recording almost nothing

  • The bug: shadow_counters showed 295/323 replays failing with HTTP 400 (signature_none_skipped: 295, last_fail_reason: "400"), so the "compression provably didn't change the agent's next action" number was computed from ~28 samples, not the live stream — the safety net was effectively off.
  • Root cause (distil/shadow.py, force_deterministic): the temp-0 replay pinned temperature = 0 unconditionally, but two API constraints reject that on the models Claude Code runs — (1) extended thinking requires temperature unset/1 (400 otherwise), and Claude Code enables thinking by default; (2) Opus 4.7+ removed temperature/top_p/top_k entirely (any value 400s), so the client omits it. Injecting temperature: 0 therefore 400'd ~every sampled request.
  • The fix: only pin an existing temperature, and never when thinking is on; otherwise replay the request exactly as sent (already API-valid). Greedy determinism is kept where the knob still exists; elsewhere the existing A/A baseline (aa_agreement / adjusted_rate) absorbs the residual sampling noise. force_deterministic is the only code in the request path that touches sampling params, and only the shadow worker calls it — so this never affected live requests, only the background measurement.

The compression mode flipped digest↔lossless-only between launches

  • The bug: the same machine sometimes ran the aggressive digest compressor and sometimes near-passthrough lossless-only — visible as wildly inconsistent per-request savings — depending only on whether ANTHROPIC_API_KEY happened to be exported when the proxy started.
  • Root cause (distil/doctor.py, subscription_mode): it classified any environment with ANTHROPIC_API_KEY set as metered → digest, even for a Claude Pro/Max user whose Claude Code traffic authenticates with the OAuth token, not the key. A volatile env var was the deciding signal.
  • The fix: the stable OAuth-login signal (~/.claude.json has oauthAccount) now wins — an OAuth login classifies as subscription (lossless-only) even with a key in the env. It fails safe: misreading subscription traffic as metered would apply lossy digest to it (the exact harm the mode gate prevents), while the reverse only leaves savings on the table. A bare key with no OAuth login is still metered; DISTIL_SUBSCRIPTION=0 forces metered under an OAuth login.

Also shipped (deploy-tier — not in the wheel)

  • The community live counter was showing a 1.44B ghost and a phantom "874.9M/day" from a single idle machine. Fixed by making every downstream layer faithful transport of the client's current emit: the worker heartbeat store is now last-write-wins by ts, the rollup community total is Σ latest-per-install (not a peak-banking ratchet), and the adoption page drops the max-anchor and the rate×86400 projection (packaging/census-worker, scripts/census_rollup.py, docs/adoption.html).

Gates

  • Full suite green; added tests for both fixes (shadow: thinking replays left valid, no temperature injected when absent; subscription: OAuth wins over a stray key, override still honored). Pinned ruff + mypy clean; distil verify / bench / validate unaffected.

1.27.0 — the coverage gate enforces the floor it advertises

Bug fix + test-debt paydown (issue #32). The coverage CI job reported success while its own log printed FAIL … Total coverage: 94.90%. Root cause: [tool.coverage.report] set no precision, and coverage.py defaults it to 0 — which rounds the number used for the --cov-fail-under decision, not just the printed report. 94.90% rounded to 95, so 95 >= 95 passed a floor the suite was actually under. For a project whose whole pitch is certified gates, a gate that lies is worse than a red build.

  • precision = 2 in [tool.coverage.report] (pyproject.toml): the fail_under comparison now uses the real two-decimal figure. (Diagnosed in #32 by @dshakes.)
  • Real coverage lifted honestly above the floor, not by lowering the bar. Meaningful tests for genuinely-untested branches: the phase-2 learned-salience trio query_flywheel / query_train / query_assoc went 80.6% → 100% (deterministic-sampling boundaries, both certify-gate rejection floors, every fail-open / skip / malformed-row path), and census.py's accrual persistence gained corrupt-channel-repair + write-fail-open tests. Total coverage now 95.48%, enforced honestly.

Gates

  • Full suite green (1818 tests); coverage ≥95% enforced at 2-decimal precision; ruff + mypy clean; distil verify / bench / validate unaffected and PASS.

1.26.0 — the community counter is monotonic; distil default --always-on

The live counter never un-counts again — count-time delta calibration

  • The bug: the community "tokens saved" odometer visibly shrank. Root cause: every census (and every heartbeat) multiplied the entire lifetime cumulative by the current calibration factor (round(lifetime × f) in build_payload, _by_model, and _current_saved_tokens). Calibration is bidirectional — as an install gathered more real usage.* samples and the factor drifted down toward a better estimate, the reported total dropped. A single active machine recalibrating downward dragged the public counter backward: tokens already saved got un-counted.
  • The fix (distil/census.py): the total is now a monotonic cumulative built from count-time-calibrated deltas — each census banks Δraw × factor-known-now and freezes it; a later factor move never restates a past increment. Monotonic by construction, and still honest: every increment is valued at the best estimate available when it was earned (exactly like invoicing each period as it closes), so the census never reports more than the provider would bill. State persists in ~/.distil/census-savings.json, is wiped with ~/.distil and on census off, and is advanced only on a real send — distil census show and every preview leave it untouched. The daily census and the near-real-time heartbeat now read the same shared total, so the live number and the board agree and neither can shrink.
  • Aggregator belt-and-suspenders (scripts/census_rollup.py): the community total and sparkline are rebuilt from cumulative positive per-install deltas (mirroring the existing rate_per_sec dtok >= 0 primitive), so even the residual pre-fix rows already in census.jsonl can't pull the public number down during the upgrade transition. For a fixed (monotonic) client this telescopes to its latest value — nothing is lost.

distil default --always-on — persist the base URL for every Claude Code launch (thanks @tolgatuncoglu!)

  • Contributed by Tolga Tuncoglu (#31): distil default --always-on writes ANTHROPIC_BASE_URL into ~/.claude/settings.json so Claude Code routes through distil on every launch without a wrapper (distil/setup.py, distil/cli.py).
  • Safety guard added on merge (distil/doctor.py): a stale or dead ANTHROPIC_BASE_URL in settings.json now fails loud in distil doctor instead of silently breaking every Claude Code session with a connection refused. distil default writes are isolated from the developer's real ~/.claude/settings.json under test.

Gates

  • Full suite green; distil bench / verify / validate PASS (byte-reversible, decision-equivalent, 60/60 adversarial). Coverage ≥95%, ruff + mypy clean.

1.25.1 — live counter: hot-swap sessions now pulse

Bug fix. The near-real-time community counter read "no machines active" even while installs were running distil. Root cause: the in-session liveness heartbeat was wired only into the legacy in-thread serve() path (_start_heartbeat_timer). Production now runs the hot-swap supervisor + worker architecture — the worker subprocess never touches census, and the supervisor (the one process that outlives every worker swap) sent no beat. A long-lived distil wrap session therefore emitted zero heartbeats until exit, so /v1/live reported active: 0 the whole time it ran.

Fix: the supervisor's _watch() poll loop (already ticking every 30 s) now pulses census.maybe_heartbeat() beside its crash breadcrumb (distil/hotswap.py). census self-throttles to ≤1/5 min and only sends when opted-in with saved tokens, so the added call is a cheap no-op most ticks. Liveness-only, honesty preserved: rate stays 0 when tokens are flat or refined downward — the odometer never projects phantom growth; only active reflects the running install. Fail-open (a census fault never touches the serving path). Verified live: active 0→1 through the new path, rate held at 0 on a downward calibration refinement.

Note: the supervisor cannot hot-reload itself (that is why workers are separate subprocesses), so a session already running when you upgrade picks this up on its next distil wrap, not mid-session.

1.25.0 — the compression frontier: nested JSON, constant-column collapse, multi-language code

Three new reversible, gate-certified techniques that close the compression gap with lossy "smart crushers" — every one keeps distil's per-request decision-equivalence proof and pure-Python install.

Nested-record JSON fold — the shape agents actually traffic in

  • fold_records (distil/compress/structured.py): the strict columnar fold only folded flat scalar records; real tool output (API responses, search hits, DB rows) nests dicts/lists and fell through to the generic digest, saving far less. fold_records extends the columnar fold to nested records — non-scalar cells render as compact JSON, the header marks which columns are JSON-encoded — for 42% fewer tokens on nested tool output. Reversible (byte-exact original one expand(handle) away), DECISION-marker-safe, reject-if-not-smaller. Wired into the tier1 path and the anthropic lossless path.

Constant-column collapse — Parquet-style entropy coding

  • A nested column repeating one value in every row (status:"active", a region, a flag) is hoisted into a single «=name<TAB>value directive and dropped from the body — the same constant-encoding a columnar database applies to a low-entropy column. Takes the nested fold from ~42% to ~62% fewer tokens on enum-heavy output. The shared value is stated once, verbatim, so the view stays maximally readable; dictionary-indexing varying columns is deliberately declined (the integer→value indirection is a decision-equivalence risk a constant hoist doesn't carry).

Multi-language code compressor — zero native dependency

  • generic_code_skeleton (distil/skeleton.py): distil's Python-only ast skeleton is joined by a language-agnostic brace-block skeleton for JS/TS/Go/Rust/Java/C/C++/Swift/Kotlin — keep every signature and brace, elide the pure-body runs between them, driven by a string- and comment-aware brace-depth scanner (braces inside strings, //, #, /* */ don't count; unbalanced/mid-edit source bails intact — save less, never corrupt). ~44% on a TypeScript file with every signature still visible. A competitor's CodeCompressor capability delivered with zero native deps — a tree-sitter grammar would be a mandatory native extension; distil stays pure-Python.
  • Now wired into the live tier1 path (was only in the offline conformal/gate path) and into smart_digest, on the active-recovery path only — an elided body needs the distil_expand handle, so it never runs on the flat-rate lossless path where the folds stay information-complete.

Live counter — liveness beat so active reflects real usage

  • maybe_heartbeat (distil/census.py) now beats every interval from any install that has saved tokens — not only when the saved-token total is climbing. A downward calibration refinement (the estimate getting more accurate, as happened at 1.24→1.25) was making a live install read as inactive; now active counts installs that are running distil. rate still reflects genuine growth (0 when flat or refined down), so the odometer never projects phantom tokens — an idle community shows an exact, unmoving total by design.

Gates

  • distil bench / verify / validate all PASS on every change (byte-reversible, decision-equivalent with recovery, 60/60 adversarial). Coverage ≥95%, ruff + mypy clean.

1.24.0 — census schema 4: the trust number, session modes, next-gen live board

The metric that drives belief — decision-equivalence, as a community number

  • equivalence {pct, shadowed} (distil/census.py): the noise-adjusted live decision-equivalence from the shadow ledger — compression provably didn't change the agent's next action, distil's core claim. pct is None until an A/A self-agreement baseline exists (never a verdict on sampling nondeterminism). The rollup aggregates it as a shadowed-count-weighted mean and the adoption page renders it as a glowing trust ring.
  • modes {interactive, headless, sdk}: session kind from the shape of the wrapped command — an agent binary (interactive) vs -p/--print (headless) vs a non-agent argv0 driving the Agent SDK (sdk). Flag presence only; the prompt/args are never read. Answers "Claude Code TUI vs claude -p vs Agent SDK" content-free.
  • API keys remain untracked, by design — billing is the only cohort split; keys are never persisted, hashed, or derived from.

Real derived metrics + next-gen live adoption page

  • Rollup now emits measured savings.rate_per_sec (Δtokens/Δt between consecutive censuses, resets dropped), as_of_ts, total_runs, avg_per_run, and a history[] community-total time series — all measured, none estimated.
  • docs/adoption.html rebuilt: a live odometer ticking at the measured savings rate (labeled projection, exact anchor, snaps on each new census), the decision-equivalence trust ring, a savings-rate readout, a tokens sparkline, a session-mode panel, and auto-refresh every 45s. Audit trail intact — every number traces to census.jsonl.
  • Both validators accept schemas 1–4 (all-or-nothing keys, prior-schema rules still enforced). Live-verified end to end: real payload equivalence {100%, 582}, modes {interactive: 28, sdk: 1}; the worker rejects out-of-range pct and unknown modes.
  • The rollup carries forward per-install dimensions so a mixed-version fleet can't blank them — a newer lower-schema ping (a v1.23 client, no equivalence) arriving after a v1.24 ping no longer erases that install's trust number.

Near-real-time community counter — opt-in heartbeat + edge aggregate

  • Heartbeat (distil/census.py): the daily census stays the exact auditable archive; a tiny content-free {v, id, tokens, rate, ts} beat — at most every 5 min and only when saved tokens grew (idle machines send nothing) — drives the live counter. Same opt-in + DO_NOT_TRACK gates, fail-open, sent from the wrap/proxy exit and a lightweight in-session timer.
  • Worker (packaging/census-worker/): /v1/beat validates + upserts latest-per-install into Upstash Redis (no history, no IPs); /v1/live sums it on read (exact total, no drift) and reports active installs + their combined rate. Both degrade gracefully without Upstash.
  • Adoption page projects the odometer forward at the active-only rate, bounded — it ticks while the community works and goes static the instant everyone idles, never inventing growth; anchors to max(live, census) so it never regresses below the archive and falls back to the census exact total when the live store is empty.

Real-time LOCAL dashboard + npm/JS-TS bridge

  • distil dashboard --web (distil/webdash.py): a localhost page fed by your own live ledger — the odometer rolls up the instant a real request books more saved tokens. Content-free, local-only, zero-dep.
  • npm distil-llm (packaging/npm/): the JS/TS bridge — npx distil-llm wrap -- <agent> resolves a Python runner (uvx/pipx/pip) so JS/TS devs use distil without touching pip; distilBaseURL() helpers point any SDK at the proxy. Closes the biggest distribution gap.

1.23.0 — census schema 3: integration attribution (SDK / headless surfaces)

Are SDKs & headless clients actually used? — now answered, content-free

  • Integration-surface + API-shape counters (distil/surfaces.py): the proxy counts each compressible request by the door it came through (wrap / proxy / gateway, from DISTIL_SURFACE set by the launching CLI command) and by API wire format (anthropic / openai-chat / openai-responses / gemini, from the request path). In-memory, flock-merged snapshot in ~/.distil/surfaces.json, flushed in serve()'s teardown before the census reads it. No key-, token-, or identity-derived data — allowlisted keys only, fail-open by contract.
  • Census schema 3 adds surfaces and shapes (request-count maps). Both validators (worker JS + CI Python) accept schemas 1–3 with all-or-nothing keys and prior-schema rules still enforced; the rollup publishes usage.surfaces / usage.shapes; the adoption page gains an Integration surface · API shape panel. Deliberately NOT tracked: API keys — never persisted, hashed, or derived from (the billing field is the closest cohort split, by design).
  • Live-verified end to end: standalone proxy records both shapes, a genuine wrapped claude -p records {wrap:1, anthropic:1}, the v3 census flows client → worker → CI → metrics branch → rollup → live panel; hostile keys (surfaces:{botnet:1}) rejected at the worker.

1.22.0 — census schema 2: usage dimensions, honest downloads, live adoption dashboard

Census (schema 2 — TELEMETRY.md updated in lockstep with the frozen-schema tests)

  • Usage dimensions: by_model (calibrated tokens saved per model id, top 5), billing (subscription|metered), agents (allowlist-only — claude/codex/gemini/aider, anything else collapses to "other" so an exotic argv can never leak). Both validators (worker JS + CI Python) accept schema 1 and 2, all-or-nothing keys, live-verified accept/reject including skew and injection attempts.
  • Honest dollars, nobody excluded: totals are calibration-corrected (never more than billed), and the rollup buckets dollars by billing — metered = real community $, subscription = notional API-rate value, published separately and labeled. Validated on a real ledger: 1.02B tokens / 5,743 runs rolled up with subscription $ correctly bucketed notional.

Adoption surfaces

  • Bot-filtered downloads: no-OS PyPI downloads are scanners/crawlers; the snapshot now records the real-vs-bot split plus per-OS and per-Python breakdowns (real-OS ≈ 1.9k/mo vs ≈ 10k bot traffic at ship time). New downloads-real badge replaces the raw count in the README.
  • Live adoption dashboard (docs/adoption.html): stat tiles with count-up, real-vs-bot split band, OS/Python/version/model bars, billing+agent chips, animated pipeline architecture diagram with live install count; linked from every sidebar, the landing strip, and the README badges. Near-real-time: census ingest re-rolls aggregates+badges on every ping; badges re-poll in 5 min.
  • Ops fix: a Vercel git-integration connected to this repo was clobbering the census worker's production deployment with function-less repo-root builds (the /v1/ping 404 outage); a root vercel.json disables git deployments — the worker deploys explicitly from packaging/census-worker/ (which gains a package.json the zero-config build needs).

1.21.0 — adoption picture: opt-in census, passive registry pipeline, live badges

Adoption & community savings (ADR 0002, TELEMETRY.md)

  • Opt-in content-free census (distil census on|off|status|show) — the answer to "how many active installs, which versions, how much is the community saving" that keeps "nothing leaves your machine" honest: OFF until explicit consent (--yes never consents; onboard asks exactly once), random install id deleted on revoke, DO_NOT_TRACK/DISTIL_NO_TELEMETRY beat stored consent, ≤1 numbers-only JSON per 24h from the proxy-exit flush, fail-open. Schema frozen by test — widening it must edit TELEMETRY.md and the test together.
  • Auditable ingest pipeline, live — zero-dep worker at distil-census.vercel.app (strict validation, 1 KB cap, numeric skew ceilings, stores nothing, no IPs) → repository_dispatch → CI re-validates (defense in depth) → appends to the public metrics branch → nightly rollup dedupes latest-per-install-id into aggregates.json + shields badges. The datastore is a git branch anyone can read.
  • Passive registry snapshot (scripts/adoption_snapshot.py, nightly) — PyPI downloads, GitHub stars/clones/views (the 14-day rolling traffic window becomes history), Docker pulls; per-source degradation; real UA + retry for shared-runner IPs; optional TRAFFIC_TOKEN (fine-grained PAT, Administration: read) for the traffic API.
  • Live displays — README badges (PyPI downloads, community tokens saved, active installs 30d) and a live adoption strip on the docs site, both fed by the metrics branch.

UX

  • Once-a-day update notice on distil wrap/distil proxy (distil/updatecheck.py): background PyPI check, one stderr line when behind, DISTIL_NO_UPDATE_CHECK=1 opts out, disclosed in TELEMETRY.md.

1.20.2 — bypass tripwire trusts the traffic marker; headless-agent examples

Fixed

  • False "no requests flowed through distil" warning on short wrap sessions. The post-run bypass tripwire checked the savings ledger, which books no rows for a session that saves 0 tokens — so a quick distil wrap -- claude -p "…" warned that the agent "may have stopped honoring ANTHROPIC_BASE_URL" even when its traffic demonstrably flowed through the proxy. The tripwire now reads the session traffic marker (written 0 at wrap start, flipped to 1 by the first proxied request — the signal built for bypass detection); it still fires on genuine bypass and stays silent when no marker exists (distil/cli.py, tests in tests/test_wrap_presets.py). Verified live: headless claude -p and the Claude Agent SDK both route through wrap with no false warning, and Claude Code 2.1.215 honors ANTHROPIC_BASE_URL on both API-key and OAuth auth (captured HEAD / preflight + POST /v1/messages?beta=true).

Docs & examples

  • Headless-agent coverage, all live-verified: new examples/python_claude_agent_sdk.py (Claude Agent SDK → bundled CLI → ANTHROPIC_BASE_URL) and examples/js_anthropic.ts (Anthropic TypeScript SDK baseURL); a "Headless agents (Agent SDK, claude -p, CI)" section in examples/README.md; headless + TS rows in the README integration table and docs/integrations.html; the previously unlisted Gemini example added to the examples table.

1.20.1 — proxy worker survives broken client writes

Fixed

  • proxy worker died (exit=-13) on long sessions. main() restores SIGPIPE to SIG_DFL (correct for CLI filters piped to head), but the proxy worker and the in-thread proxy are long-lived network servers: a write to a client/upstream socket that had hung up killed the whole worker with signal 13 instead of raising a catchable BrokenPipeError, dropping in-flight state on each respawning. Both serving paths now reinstate the server-safe SIG_IGN (distil/hotswap.py, distil/proxy.py), next to where they already override SIGINT for the same "server, not a filter" reason — a broken client now aborts only that one request. Regression test asserts a worker survives 30 abrupt RST disconnects and still serves (tests/test_hotswap.py), verified to fail without the fix.

1.20.0 — proof surfaces: live gate, OpenAI/Gemini parity, gateway keys, encrypt-at-rest

Certification & honesty

  • Nightly live gate (.github/workflows/live-cert.yml) — re-certifies the whole corpus against a real model on a nightly cron: ONE pooled TOST over every turn (distil certify -t corpus/), majority-of-3 sampling (--samples), and an A/A self-agreement control on by default (--no-aa-control opts out) so a turn where the model disagrees with itself cannot indict compression — the shadow v2→v3 lesson applied to the certify gate. Budget-capped with a hard --max-live-calls ceiling so an unattended run can never spend silently; --margin 0.10 is the live regression margin, distinct from the offline 0.02 proof margin. Live-validated before shipping (claude-haiku-4-5, 28 pooled turns: mean diff −0.036, p=0.0415, PASS). Per-commit gates remain the synthetic offline oracle — the A/A control is a provable no-op there; both layers are labeled precisely.

Adapters

  • First-class OpenAI adapter (distil/adapters/openai.py) — Chat Completions and Responses API shapes, recency carve-out, and the same Tier-0/1 machinery as the Anthropic adapter — including expand-tool injection and output shaping for the Responses shape (bounded loop, same PAYG gating as the messages path) and query-aware intent from input_text + function_call args. Live-unverified against a real OpenAI endpoint (offline shape coverage only).
  • Gemini parity (distil/adapters/gemini.py) — recency carve-out, query-aware intent from functionCall args, output shaping via shape="gemini", and expand-tool injection as a functionDeclarations entry with a bounded functionCall→functionResponse loop. cachedContent needs no guard: cached turns live server-side and never appear in contents. Live-unverified against a real Gemini endpoint (offline shape coverage only).

UX

  • Per-agent wrap presets (distil/onboard.py:AGENT_PRESETS) — distil wrap -- claude|codex|gemini|aider auto-selects the correct env var (ANTHROPIC_BASE_URL / OPENAI_BASE_URL / GOOGLE_GEMINI_BASE_URL) and upstream; cursor-agent omitted (env var undocumented). Explicit --env-var/--upstream always win. Prints preset: <label> detected → <VAR>.
  • Proof Ledger (distil/proof_ledger.py) — end-of-session printout on distil wrap exit: calibrated tokens/cost, shadow verdict with honest suppression labels, restorability. Silent on zero requests. Opt-out: DISTIL_NO_LEDGER=1.
  • Zero-traffic tripwire — distil wrap warns when a known-agent session ends with 0 proxied requests (upstream env-var contract broken, likely agent update). Tests pinned in tests/test_upstream_contracts.py.

Security

  • Encrypt digest originals at rest (distil/atrest.py) — HMAC-SHA256-CTR + encrypt-then-MAC, DSTL1 magic header, restore.key at chmod 0600. Protects against backup/sync leakage and cross-user reads on shared filesystems. Does not protect against same-UID attackers (documented in THREAT_MODEL.md). Legacy plaintext files still load transparently. Opt-out: DISTIL_NO_ENCRYPT_AT_REST=1.

Gateway

  • Gateway keys (distil/gateway_keys.py, distil gateway keys issue|list|revoke) — issued dsk- keys hashed at rest (SHA-256); auth fails closed; upstream never sees the distil key (both header carriers stripped unconditionally). --require-keys forces key auth even before any keys are issued. --tenant-rpm and --tenant-daily-tokens enforce per-tenant rate and quota limits — including with key auth off (credential-derived tenants) and per-key overrides set at issue time; every advertised control was verified enforcing under independent security review, which also judged the at-rest construction sound. Default path unchanged (anonymous hash-based tenant IDs still work without keys).

Benchmarks

  • OTel session correlation (distil/otel.py) — distil.session.id span attribute on every proxied request, enabling per-session trace correlation in any OTel backend.
  • Referee scorecard (benchmarks/scorecard.py) — grades any compressor on distil's five invariants; used by the head-to-head harness and linkable from docs/benchmarks.html.

1.19.0 — distil validate: adversarial real-path gate

Added

  • distil validate — a validation harness that drives the real compressor against a battery of diverse and hostile inputs (huge / unicode / deeply-nested / malformed tool output, marker-injection that mimics distil's own << handle >> stubs, secret-looking strings) and asserts the load-bearing guarantees on every one: reversibility (every handle recovers its exact bytes), reject-if-bigger, recency-exact (latest tool output byte-identical), fail-open (no input makes the compressor raise), and content-free (content reaches only the local restore store, never a telemetry file — proven with a unique marker). Runs in CI on every push alongside bench and verify; exits non-zero on any violation.

This is the adversarial layer between "green unit suite" and "GA-ready" — it exists because a passing suite kept coexisting with real-traffic bugs in code paths the corpus never exercised.

1.18.2 — calibration robustness (capture-miss immunity)

Fixed

  • Calibration is now immune to token-usage capture misses. On real Claude Code traffic some requests logged the new (uncached) input but not the cached prefix, giving billed ≪ est and a ~0.001 ratio that polluted the store. Now record() filters any pair outside the plausible band [0.5, 3.0], and factor() computes the median of only in-band ratios — so it is robust to garbage already written by an older or concurrent producer. A public headline can never be multiplied by a capture artifact. Measured correction on real traffic is ~1.03–1.05 (the heuristic is accurate because the large cached system prompt dominates and is well estimated) — smaller than the 15–20% feared.

1.18.1 — calibration: prompt-cache correctness + sanity bounds

Fixed

  • Calibration now accounts for prompt caching. usage.input_tokens counts only the uncached tokens; the cached prefix is billed under cache_read/cache_creation. Comparing the heuristic's full-request estimate to input_tokens alone made the factor collapse (~0.001 on real cached traffic). scan_usage now captures the cache fields and calibration uses the full billed input (input + cache_read + cache_creation). dissect.calibration() had the same bug — also fixed.
  • Sanity-bounded factor. A correction outside [0.5, 3.0] is a data problem, not a tokenizer difference — the factor falls back to identity, so bad data can never poison a public headline.

1.18.0 — self-calibrating token counts (billing-grade, no network)

The offline heuristic is ~15–20% off the real BPE (40%+ on dense code — measured). But distil is a proxy: it sees the provider's real usage.* on every response. It now learns the correction from that pairing and reports token counts that converge to the real tokenizer — with no per-string network call. The compression percentage was always exact (numerator and denominator use the same estimator); this fixes the absolute counts the leaderboard shows.

Added

  • Self-calibrating token counts (distil/calibration.py, mechanisms A + C + D). A per-model store of (heuristic estimate, billed) pairs — integers only, never text — recorded by the proxy from usage.input_tokens. factor() returns the correction (aggregate ratio blended with the per-request median, robust to outliers) and is identity (1.0) until 20 observations, so an uncalibrated install reports exactly today's numbers — no regression, no early skew. The leaderboard (text + HTML) applies it to the absolute totals and prints "calibrated to your billed usage (N requests, ±X%)"; the percentage is unchanged (scale-invariant).
  • --tokenizer subword (mechanism B) — a length-aware offline BPE approximation, still zero-dependency: it charges longer identifiers more (as BPE does) and adds a surcharge for multi-byte characters. Measured closer to the real count than the flat heuristic (33% vs 41% error on a code sample); a better base for calibration to correct the rest of.

Validated: 13 unit tests (convergence, identity-until-proven, content-free, per-model + pooled, CI, bounded reservoir, corrupt-store-safe, subword properties); a live Anthropic count_tokens comparison; mypy clean; full gate + distil bench + distil verify PASS.

1.17.0 — semantic bridge (always-on query↔answer matching, zero-dep)

Phase 2 (1.16.0) pinned semantically-relevant lines only near a lexical hit and only once a model was trained. This closes that gap: an always-on, zero-dependency semantic bridge lets a query term match an answer term that shares no spelling — with no model and no embeddings.

Added

  • Semantic bridge (compress/lexicon.py) — four composable, pure-Python mechanisms unioned into the query-relevant keep set, all additive (they only ever widen keeps, so a wrong match wastes a little compression, never drops an answer):
    • ① a suffix-stripping stemmer + a curated technical synonym map (retry↔attempt, limit↔max/cap/threshold, timeout↔deadline/ttl…), with compound-identifier splitting so max_attempts → {max, attempts} bridges to {limit, retry};
    • ③ char-trigram Jaccard for typos / near-morphology (config↔configs);
    • ④ optional distributional vectors — pure-Python cosine over a bundled table, no torch/numpy at runtime; inert (no-op) until a table is provided, so the zero-dependency posture is preserved.
  • Flywheel-learned associations (query_assoc.py) — the moat: distil learns your vocabulary (tenant↔org, max_tries↔retry) from real expands, by expand-conditioned co-occurrence. Content-free — only hashed term pairs are stored, joined to the existing content-free expand log. Rebuilt by distil query-relevance.

The bridge is on by default in the digest (phase 1 + bridge); the phase-2 learned model adds proximity on top. Validated: 13 unit tests (each mechanism + additive-safety + content-free); a claude -p-confirmed answer (timeout→deadline_ms) recovered by the bridge that lexical misses; distil bench (verdict-retention) + distil verify (byte-fidelity) PASS.

1.16.0 — query-aware salience, phase 2 (learned semantic relevance)

Phase 1 (1.15.0) pins tool-output lines that lexically match the agent's intent — a grep hit, a config key, a SHA. Phase 2 adds the semantic case a fixed lexical rule can't reach: ask "what's the retry limit?" and the answer line max_attempts = 5 shares no token with the query, so phase 1 folds it (recoverable, but a round-trip). Phase 2 pins it inline.

Added

  • Learned query-relevance scorer (distil/compress/query_relevance.py) — an embedding-free, stdlib-only logistic model over query-conditioned features (lexical overlap, selectivity, proximity to a lexical hit). It's layered additively over phase 1: its kept lines are unioned in, so it can only ever widen the keep set — reversibility and the decision- equivalence certificate are untouched. Gated on promoted weights: with none, behavior is exactly phase 1, so shipping it is behavior-neutral.
  • Content-free expand flywheel (distil/query_flywheel.py) — the model's training labels come from distil's own traffic: which digested block the agent expanded, paired with the query live at digest time. Records only numeric feature vectors + the expand outcome — never a raw prompt, response, tool result, or query term. Off by default; the live proxy enables it under --expand. Sampled and fail-open.
  • Train + certify (distil/query_train.py, distil query-relevance) — trains on the flywheel labels (reusing the keep-model's logistic trainer, generalized to any feature width) and promotes weights only if held-out recall beats the phase-1 lexical baseline at a precision floor. Additive-only makes it decision-safe by construction; the gate guards compression waste.

Validated: recall 1.0 vs a 0.0 lexical baseline on the semantic case; a claude -p-confirmed answer line that phase 1 folds is recovered by phase 2; distil bench + distil verify PASS.

1.15.4 — GA hardening: structured-audit fixes

A three-front structured audit (proxy request path, compression correctness, policy/telemetry privacy). The content-free telemetry guarantee, the opt-in transcript correlation, and the #28 fix were all independently verified clean. Six real findings fixed, each with a proof test; the async silent-drop fix from the prior batch is included here.

Fixed

  • [HIGH] Subscription-safe default now applies to direct distil wrap/distil proxy. Only the managed distil default install auto-selected lossless-only before; a bare distil wrap -- claude on a subscription ran the Tier-1 digest with no expand tool — leaving irreversibly-lossy stubs in a subscription session. Now a detected subscription with no explicit mode flag defaults to lossless-only. An explicit --expand/--verbatim/--lossless-only always wins.
  • [HIGH] Recency exemption was silently bypassed for the standard tool_result list shape. The list-content path dropped the is_recent flag (the string and OpenAI paths passed it), so the agent's most-recent output — which the real Anthropic SDK always sends as a list — got folded instead of kept byte-exact. A regression introduced with the 1.15.1 columnar fold.
  • [MED] --lossless-only --shape-output X no longer falsely claims to shape. It printed "output shaping: X" at startup while the handler correctly suppressed shaping on lossless-only. Now it warns the request is suppressed (sync + async proxies).
  • [MED] distil proxy --async no longer silently drops --expand/--session-delta/--shadow (from the prior batch) — it names the ignored flags and points at the standard proxy.
  • [LOW] Disk restore now has the in-memory collision guard — a 32-bit handle collision across sessions can no longer clobber and expand to the wrong bytes.
  • [LOW] fold bails when a JSON key contains , (the column-header delimiter) instead of emitting a mis-keyed table.

1.15.3 — honor explicit --expand on a subscription (#28)

Fixed

  • --expand is no longer silently disabled on a subscription (#28, thanks @pliablepixels). Subscription sessions run lossless-only, which forces verbatim and turns the Tier-1 digest off — so --expand did nothing there. The verbatim force exists because an unrecoverable stub is irreversibly lossy; but --expand injects distil_expand, which makes every stub recoverable, so that hazard doesn't apply. An explicit --expand now lifts the force even on a subscription: maximum recoverable compression, nothing irreversibly lost, with a one-time startup notice. The default is unchanged — no --expand still means lossless-only. Genuinely-lossy output shaping stays metered-only (it rewrites the response, which expand can't recover).

1.15.2 — CI portability + docs

Maintenance release. No shipped-code behaviour change from 1.15.1 — the package is byte-identical apart from the version; this cuts a clean tag over the test/docs fixes below.

Fixed

  • CI green on Windows and under load — two portability bugs in the adopted dissect tests (both passed on Linux/macOS): a report file was read with the platform default codec (UnicodeDecodeError on Windows cp1252 for the « fold marker), and the streamed detail record was read before its line flushed (IndexError under CI timing). Test-only.

Docs

  • The landing page, getting-started FAQ, and techniques page now depict the full 1.15 line — content-type keep policy, query-aware salience, columnar fold, and distil dissect. The "will it save me money?" answer no longer under-sells subscriptions: lossless mode does cut tokens per turn (headroom + rate-limit room), it just doesn't lower a flat-rate bill.

1.15.1 — distil dissect + lossless columnar fold for subscriptions

Driven by the same independent power-user (@pliablepixels, #24 / #26 / PR #27). Both a big new observability feature and a real subscription-savings gap addressed.

Added

  • distil dissect (#26, PR #27, thanks @pliablepixels) — a per-session deep-dive report: savings by model and by mechanism (digest vs cache-delta), the digest inventory (blocks by kind, largest folds, re-fold churn, restore recoverability), billed usage captured from API responses with a heuristic-calibration figure, latency by path (the --expand buffering tax as a measured number), quality loops, and a "worth your attention" anomaly list that auto-detects the #25 signatures so that class of silent failure can't hide again. Optional, strictly opt-in transcript correlation names tools/files/prompts; everything else stays content-free. New content-free session logging on the proxy is fail-open and off the request path. distil dissect [session] [--html|--json|--serve].
  • Lossless columnar fold on the subscription/lossless path (#24) — a JSON array of flat records (extremely common tool output) now folds to a compact, self-describing table (all rows inline, no recovery handle to invite an unavailable distil_expand), ~70–79% lossless and ToS-safe. Recent tool_results stay byte-exact (never fold) so the agent's latest output is unchanged. Inherits fold's decision-equivalence certification.

Fixed

  • Subscription onboard clarity (#24) — the note now states plainly that lossless mode does cut tokens (JSON minify + run collapse + columnar fold), it just doesn't lower a flat-rate bill.
  • Adopted with review fixes over the contributor's branch: argv persisted as command[:1] only (no credential-in-flag leak to the session manifest), and fcntl.flock on the per-request JSONL append (concurrent-write safety).

1.15.0 — query-aware salience, content-type keep policy, expand reliability

The digest gets genuinely content- and intent-aware, and two reliability bugs on the subscription path are fixed. Several items began as issues/a PR from an external tester (@pliablepixels, #22–#25) and are credited on the commits.

Added

  • Query-aware salience — distil is a proxy, so at compress time it holds the agent's intent (its tool_use arguments + latest ask) in the same request as the output being compressed. Lines that match a discriminating intent term are now additively pinned, so the one line the agent is looking for survives even in arbitrary output where no fixed rule knows which line matters (a grep hit, a config value, a SHA). No post-hoc compressor has that query/output pairing. Strictly additive: the keep set only widens, so reversibility is untouched and the certificate can only hold or improve. A non-discriminating term (one matching most lines) is dropped, preserving compression. distil/compress/intent.py; spec in specs/query-aware-salience.md. Example: a 4,152-token log with the answer buried in 600 neutral lines → ~163 tokens with the answer kept.
  • Per-content-type keep policy (#23, thanks @pliablepixels) — a new distil/compress/keep_policy.py classifies a block (log / traceback / diff / generic) and keeps each kind's load-bearing lines: a log's pass/fail verdict, a traceback's stack frames, a diff's hunk headers — on top of the generic error/DECISION net. Supersedes the 1.14.x inline verdict rules with a cleaner, extensible module (the "per-content-type codec" the tier-1 docstring long promised); the outcome-aware dedup layer rides on top unchanged.
  • Shadow observability (#25, thanks @pliablepixels) — shadow sampling now keeps content-free counters (requests seen / sampled / replays attempted / failed + last reason / recorded). distil shadow-stats explains a 0 recorded result ("19 seen, 2 sampled, 2 replays failed (last: 401)") instead of a silent "no samples yet", so a failing replay path (e.g. OAuth rejecting proxy replays) is visible instead of indistinguishable from bad luck.

Fixed

  • distil_expand tool_use escaping on the streamed path (#25, thanks @pliablepixels) — the expand gate keyed on handles created this request, but RestoreStore persists to disk, so a streamed turn that digested nothing new yet referenced an older stub emitted a distil_expand call with no tool injected and no expand loop — and it escaped to the client as "No such tool available". The gate now keys on any recoverable handle in the outgoing conversation, injecting the tool and buffering to resolve the call server-side.

Changed

  • Subscription onboarding clarity (#24, thanks @pliablepixels) — when a flat-rate subscription is detected, distil onboard now states plainly that lossless mode trims context + latency but does not reduce token count or cost. (The suggested lossy default for subscriptions was declined: it is not provider-ToS-safe and --session-delta needs the distil_expand tool that lossless mode intentionally withholds.)

1.14.1 — outcome-aware routing + verdict gate

Added

  • Outcome-aware routing (tier-1) — the first content-type profile. When the log's own verdict says GREEN (tests passed / build succeeded, nothing failed), ERROR/WARN stdout is by definition noise the SUT logged on purpose: dedup tightens to one sample per shape. Red or unknown outcome keeps the cautious 2; explicit max_repeats still overrides; everything folded stays recoverable. Power-user log shape: 19.9k tokens → ~93 with verdict + error signal in front of the agent.
  • Verdict-retention self-check in distil bench — the gate digests a canonical green and red test log and requires the verdict line to survive both (the comparison's own substring check). A digest change that compresses away the answer flips CI red instead of shipping silently.

1.14.0 — verdict-aware digest: keep the answer, fold the noise

Driven by an independent power-user comparison on a live repo (4 tools, 3 tasks): distil tied for best on bug-fact retention and was the only tool with measured byte-exact reversibility — but on a 22,971-token passing test log it folded the one line that mattered (1955 passed) into a handle while keeping repeated ERROR/WARN stdout. Both halves of that inversion are fixed.

Added

  • Verdict preservation (tier-1) — _SUMMARY_RE in the digest keep-net pins command result lines verbatim: vitest/jest/pytest/mocha counts, cargo test result:, go ok/PASS/--- FAIL:, gradle/maven BUILD SUCCESSFUL|FAILED, and exit code N. A green run's verdict (which carries no error keyword) can no longer be compressed away.
  • Verdict preservation (salience) — SalienceKeepModel scores the same verdict lines at the 1.0 never-drop floor (they previously scored 0.3–0.6 and dropped while ERROR noise scored 0.95). Single source of truth: tier-1's _SUMMARY_RE.
  • Error-noise dedup (tier-1) — near-identical error/warn repeats (same line shape after normalizing digits/hex) keep their first 2 occurrences as signal; the rest fold behind the existing handle markers, fully recoverable. DECISION: and verdict lines are exempt.

Changed

  • On the comparison's log shape (passing suite, looped on-purpose ERROR/WARN stdout, verdict near the tail): 19.9k tokens → ~131 tokens with the verdict and first-occurrence error signal in front of the agent — previously 90% reduction but the answer required a second round-trip to recover.

Everything is additive to the keep-rules and reversible; no wire, config, or API changes.

1.13.0 — trustworthy shadow gate: deterministic decision-equivalence + mode visibility

Promoted to GA on live validation: 100% decision-equivalence over 116 sampled production requests (0 decision changes), with a temperature-0 A/A self-agreement baseline of 31/31 confirming the result is compression fidelity, not sampling noise. (Prerelease train: rc1–rc7.)

[Withdrawn 2026-09-04] This 100%/31-31 number was computed over a biased subset of traffic (thinking-heavy turns all failed to replay and were silently excluded — see 1.51.1); the current, honest reading is 44 A/B and 11 A/A samples with raw agreement 81.8% [67.3, 91.8] and self-agreement 84.8% [71.8, 92.4] over 46 byte-identical replays (p=0.63), not the number above.

Added

  • Status-line mode chip — the compression mode is now visible at a glance: ⬢ digest · ◇ lossless · ▪ verbatim, rendered right after distil. Read from the ledger's most recent row via ledger.latest_mode().
  • Docs — plain-English + technical mode explanations (digest / expand / lossless-only / verbatim), the billing→mode auto-default (distil default/onboard), corrected shadow-gate mechanics (50 A/B + 30 A/A), --shadow 1.0 high-fidelity validation, and a validated decision-equivalence callout, across README and the docs site.

Fixed

  • Shadow decision-equivalence is now measured deterministically. The A/A self-agreement baseline was reading ~38% — not because compression changed the agent's decision, but because the replay ran at the agent's live sampling temperature, so the model disagreed with itself on identical input. Both the served and replay sides of every shadow sample are now re-issued at temperature 0 (shadow.force_deterministic), never reusing the live hot response. A/A collapses toward ~100%, so the A/B rate becomes a real compression signal instead of noise. Signature methodology bumps to v3; v2 samples are scoped out (never averaged with v3), so the gate restarts on a clean baseline. This GA ships with the v3 gate passed (see the validation note above).
  • As a side effect, the streaming path no longer buffers the upstream response body for shadow (it re-issues its own calls), removing that per-request copy.

1.13.0 — 1.13.0rc6 — seamless hot-swap: upgrades apply to live sessions, no restart

Added

  • Savings ledger rows are stamped with the compression mode (verbatim / lossless-only / digest), so "why was ▼ low on this session?" is answerable directly from savings.jsonl instead of by inference — a lossless-only row saving ~0% is subscription safety working as designed; a digest row saving ~0% is genuinely low-redundancy content. Optional; pre-1.13 rows read as unknown mode.

Fixed

  • distil stats / leaderboard decision-equivalence now scopes to the current signature version and reports the noise-adjusted rate (like the status line), instead of the stale, un-adjusted v1 number.
  • Shadow decision-equivalence is trustworthy again (decision-signature v2, see docs/adr/0001-shadow-decision-signature-v2.md). The v1 signature hashed tool arguments verbatim (normalizing only Python code), so wording jitter — ls -la vs ls -la, re-serialized JSON — read as a changed decision, inflating the measured divergence for both compressed and self-replay traffic (a live ledger showed only 72.7% A/A self-agreement). v2 canonicalizes formatting whitespace on all arguments without merging genuinely different tokens. The signature algorithm is now versioned (SIG_VERSION); ledger rows are stamped with sig+build and the verdict scopes to the current algorithm (shadow-stats --all reads every row), so old-version rows can no longer drag a live verdict. The status-line verdict now requires robust evidence (≥50 A/B, ≥30 A/A samples) before showing ✓/✗, warming as de baseline N/30 otherwise. Alarm thresholds are unchanged — the sample gate, not a looser threshold, stops false alarms, so real degradation still trips it.
  • Hot-swap supervisor no longer cries wolf when a worker dies during a non-atomic reinstall. pip/uv --force-reinstall deletes the package files before rewriting them; a worker spawned in that ~1s window dies importing half-gone code, and the supervisor logged a scary WARNING proxy worker died;respawning and tight-looped. It already self-heals once the install completes — now, when installed_version() is momentarily unreadable, it logs at INFO and waits _UPGRADE_SETTLE_S before respawning. (Surfaced by reinstalling a shared pipx venv under live distil wrap sessions.)
  • Shadow health no longer shows a red ✗ degraded verdict before the A/A noise baseline exists. adjusted_rate() silently falls back to the raw, un-adjusted rate when there is no baseline, so the status line was painting sampling nondeterminism as compression harm (e.g. a scary ✗de 36.0% over 25 samples with a 3/10 baseline). The verdict glyph now gates on aa_agreement() and shows a neutral de baseline N/10 while warming; shadow-stats --json nulls its adjusted_* fields until the baseline lands instead of labelling raw values "adjusted". Display-only — never affected routing, compression, or savings.

Added

  • Seamless proxy hot-swap (POSIX, on by default): distil wrap now runs the proxy as a supervised subprocess on a wrap-owned listener FD. When pipx upgrade (or pip) puts a new version on disk, the wrap spawns a fresh worker — new code, same socket, same port — health-checks it, then drains the old one: in-flight requests (including long LLM streams) finish on the old worker while new requests land on the new one. The agent session never restarts and its ANTHROPIC_BASE_URL never changes.
    • Zero request-path overhead: supervision is out-of-band; the upgrade poll is one metadata read every 30 s in a daemon thread.
    • Fail-safe twice over: a worker that doesn't report ready is discarded and the old one keeps serving; a supervisor that can't start falls back to the historical in-thread proxy. The feature can never cost a session.
    • A worker that dies mid-session (crash/OOM) is respawned automatically — the same self-heal contract the in-thread accept loop had.
    • Manual trigger: kill -USR1 <wrap pid>. Opt out: DISTIL_HOT_SWAP=0.
    • Windows keeps the in-thread proxy and the existing skew warning (FD inheritance is POSIX-only — same accepted platform split as file locking).
  • distil proxy-worker (internal) — the supervised worker entry point.
  • distil upgrade now says which sessions hot-swap on their own instead of telling you to restart everything.

1.12.0 — soaking as 1.12.0rc4 since 2026-07-06 — statusline honesty round 3: "✓ on" means traffic actually flows

First release through the new rc + soak pipeline (runtime code → rc first).

Added (post-rc4, headed for rc5)

  • OpenTelemetry GenAI spans (opt-in): pip install 'distil-llm[otel]' emits gen_ai.* semantic-convention spans per proxied request with distil.tokens.original/compressed, distil.compression.ratio, and distil.shadow.sampled attributes. Strict no-op without the extra; an OTel failure can never break the request path. Core stays zero-dependency.
  • Supply chain: CycloneDX SBOM attached to GitHub releases; weekly OpenSSF Scorecard on main; PEP 740 Sigstore attestations confirmed active.
  • docs/EVALUATION.md: the evaluation methodology — why compression ratio without a task-success delta is meaningless, the E7 negative result, the A/A nondeterminism baseline, and what the trajectory certificate does/doesn't prove.

Fixed (post-rc4, headed for rc5)

  • CI red on main since 1.11.2 — test-side, not product: the wrap signal tests synchronized on fixed sleeps and lost the race on loaded CI runners (SIGTERM landing before the handler installs kills the wrap raw). Children now write a readiness marker after arming; tests wait on it.
  • shadow.jsonl appends now flocked like ledger.json — shadow is on-by-default since rc3 and rc4 rows exceed the size where bare appends are atomic; concurrent sessions can no longer tear rows.
  • distil doctor no longer silently drops a crashed Claude Code check — it reports FAIL like every other check.

Fixed

  • "✓ on" no longer trusts env vars alone — a wrapped agent that bypasses the proxy is called out. Found live: a claude.ai-subscription (OAuth) Claude Code session keeps DISTIL_SESSION and the loopback ANTHROPIC_BASE_URL in its env yet sends model calls straight to api.anthropic.com (verified with lsof: direct TLS to the provider, zero ledger rows, proxy healthy). The 1.11.1 honesty fix checked routing setup; this one checks routing reality. distil wrap now writes a per-session traffic marker (~/.distil/sessions/<sid>, "0" at start), the proxy's first proxied request flips it to "1", and a marker still at "0" after a 3-minute grace shows "⚠ wrapped, agent bypassing proxy" (minimal mode: "⚠ bypassed") instead of "✓ on". Markers are single-writer (no locking), best-effort (never block the wrap), swept after 7 days, and standalone distil proxy never fabricates one.

Added

  • A/A noise baseline makes the de number interpretable (rc4). Soak found raw compressed-vs-full agreement at 47% — alarming until you notice the comparator's bar is "same tool with identical normalized arguments" under live sampling: the model disagrees with itself on identical requests (a Bash command worded two ways is a "changed decision" with zero compression involvement). A third of shadow samples now replay the SAME compressed request twice, measuring self-agreement; the statusline de rate is reported relative to that baseline, and distil shadow-stats shows the full decomposition (raw / self-agreement / adjusted). Shadow rows also carry content-free evidence now — request digest + both decision signatures — so any divergence is diagnosable instead of a bare false.
  • Shadow decision-equivalence sampling is on by default (rc3): 2% of wrap requests. It was opt-in (--shadow, default 0) — so nobody ran it, the statusline's de 1/25 counter sat frozen for a week implying live measurement, and the launch gate's "✓de ≥ 99% at n ≥ 25" evidence could never accrue. Per the intelligence-is-the-default rule the flag is now an opt-out: --shadow 0 disables, --shadow 0.1 collects faster. Cost is explicit: a sampled request is re-run uncompressed for comparison, so the default adds ~2% tokens.
  • Wrap-signal breadcrumbs (rc3). Tonight's quit produced no .exit file — because the killer took out the wrap with the child (process-group kill; terminal-tab SIGHUP), so the child-exit path never ran. The wrap's SIGTERM/SIGHUP handler now appends "wrap received SIGNAME" to the .exit file before dying — the only trace a group kill leaves. Both lines together read as a story: wrap received SIGTERM; child exit 143.
  • Child-exit breadcrumb (rc2). Soak day 1 hit a recurring silent agent quit — no crash report, no error in the transcript, no way to tell an OOM abort from a clean exit after the fact. The wrap is the only witness, so it now records how the child ended (~/.distil/sessions/<sid>.exit: "exit code N" / "signal NAME" + timestamp); scripts/soak-report.sh prints it per session.
  • Terminal private-mode reset on wrap exit (rc2). A crashed TUI leaves xterm modes on that tcsetattr can't undo — mouse reporting (the 65;76;9M junk on click), bracketed paste, the alternate screen, a hidden cursor. The wrap's restore now resets them explicitly; all idempotent on clean exits.
  • Test-env hygiene: tests/conftest.py sandboxes DISTIL_HOME and strips the inherited DISTIL_SESSION for every test — dogfooding developers run the suite from wrapped terminals, and no test may touch the real ~/.distil.

1.11.4 — 2026-07-05 — Release hardening: chaos suite in CI, rc + soak policy, launch gate

No runtime code changed in this release — it hardens the process that ships the runtime. The 1.10.0→1.11.3 day (six releases, each fixing the previous, all correct in review and wrong under real use) showed the release gate was blind to signal/lifecycle chaos and had no bake time. Both gaps close here. (Per the new soak policy this release is soak-exempt: tests/CI/docs/tooling only.)

Added

  • Chaos suite in CI (tests/test_chaos.py): the ad-hoc harnesses used to verify the 1.11.3 Ctrl+C fix are now permanent, bounded tests that run on every push — a ~400-signal sustained SIGINT hammer against distil wrap with a live child (pins the 1.11.3 immune-parent property; the 1.11.2 structure fails this), and a crash-the-accept-loop test proving the wrap proxy self-heals and keeps answering on the same port (the 1.11.0 self-heal path was previously untested).
  • rc + soak release policy (RELEASING.md): any release that changes runtime behavior ships as X.Y.ZrcN first and bakes ≥ 3 days on real traffic before the final. rc tags are fully wired: GitHub release marked prerelease, PyPI gets the rc (pip ignores prereleases unless --pre), Homebrew and the Docker image skip rcs, release.sh detects rc versions and adjusts its preflight (changelog entry lives under the final).
  • Launch gate (docs/GA_READINESS.md): a binary, evidence-based checklist separating engineering GA from the marketing launch — 14 quiet days at head, external beta, live decision-equivalence at n ≥ 25 from multiple users, human fresh-install walkthrough on all three OSes, claims re-audit at the launch commit.

Fixed

  • Statusline de honesty (rc3, honesty gap #3). A sub-25 sample count now shows de n/25 only while the shadow ledger was fed within the last 24h; otherwise de idle — a frozen counter must not read as live measurement.
  • Signal-handler breadcrumb wrote an empty file (rc3). _signal_breadcrumb used time.strftime but proxy.py only imported time inside wrap_run — the NameError was swallowed by the handler's best-effort except, leaving a created-but-empty .exit. Caught by the new SIGHUP test; time is now a module-level import.
  • release.yml would have served an rc to everyone. A v*rc* tag previously bumped the Homebrew tap and pushed the Docker image as latest — both now final-only, and the GitHub release for an rc is marked prerelease.

1.11.3 — 2026-07-05 — Ctrl+C fix, take two: the wrap parent is now immune to SIGINT entirely

Fixed

  • Rapid Ctrl+C could still kill a wrapped session (escape path in the 1.11.2 fix). 1.11.2 caught the Ctrl+C KeyboardInterrupt only while blocked in proc.wait(); a press landing while the parent was executing the except clause itself (users mash Ctrl+C — Claude Code literally prompts "press ctrl-c again", and a held key auto-repeats) escaped the loop, tore the proxy down under the live agent, and killed the session on its next API call. Reproduced under a SIGINT hammer: the 1.11.2 structure died after ~1.7k signals with the agent still alive; the new one survived 1.7M. The wrap parent now installs a no-op SIGINT handler for its lifetime — immune to any number and timing of presses. A Python-level handler (unlike SIG_IGN) resets to default across exec, so the agent still receives its Ctrl+C normally (verified empirically). SIGTERM keeps terminate-child + flush-savings + exit semantics.
  • scripts/release.sh: dropped the stale distil/__init__.py version-literal check — the literal was removed when the version became single-sourced from pyproject.toml, so the check could only fail.

1.11.2 — 2026-07-05 — Ctrl+C no longer kills wrapped sessions; fresh-install statusline honesty

Fixed

  • Ctrl+C no longer tears down the proxy under a live agent. A terminal Ctrl+C is delivered to the whole foreground process group; agents like Claude Code survive the first press (it cancels the turn, not the app), but distil wrap treated it as shutdown — exiting and leaving the agent pointed at a dead port, so the session died on its next API call. The wrap now keeps waiting through SIGINT (the child owns that signal); SIGTERM keeps its terminate-child + flush-savings + exit semantics.
  • Statusline honesty for fresh installs. A routed session (distil wrap env present) with an empty ledger was told to run distil wrap -- <agent> — the exact state every new user hits first. It now shows "✓ on · no savings yet"; the wrap hint remains only for genuinely unrouted shells.
  • Alias-mode verify hint fixed. distil default told users to check echo $ANTHROPIC_BASE_URL, which is empty by design in alias mode (the URL is injected only into the wrapped agent's env). It now says type <agent> (should show the distil wrap alias); the env-var check applies to --always-on only.

1.11.1 — 2026-07-05 — Statusline honesty, pre-1.10 warning in terminal, distil reset

Added

  • distil reset — archives the savings ledger to savings.jsonl.reset-<utc> (non-destructive, auditable) and starts fresh on post-1.10 accounting; --shadow also resets decision-equivalence stats. For ledgers dominated by pre-1.10 records whose savings may be overstated.

Fixed

  • Statusline honesty: "✓ on" now means routed. The idle segment said "✓ on" even in a session whose requests went straight to the provider (no distil wrap, no loopback base URL). Unrouted sessions now show "off — session not routed".
  • Pre-1.10 overstatement warning reaches the terminal. distil stats text output now prints the legacy-accounting footnote (was HTML-only, despite the 1.10 changelog claim), with a pointer to distil reset.
  • Windows: distil default --undo test no longer assumes a service manager exists (none is wired on Windows).

1.11.0 — 2026-07-05 — Ops-ready: debuggable fail-open, crash-safe accounting, health probes; claims audit

Added

  • GET /distil/health on all three entry points (sync proxy, async proxy, gateway): unauthenticated liveness probe for load balancers / k8s readiness checks. Answers locally — never touches the billed upstream.
  • Debug escape hatch for fail-open paths. DISTIL_DEBUG=1 (or DISTIL_LOG_LEVEL=<level>) logs every swallowed compression/learning/shadow exception to stderr with a traceback, via a distil logger that never touches the root logger. Silent-by-default is unchanged.
  • Restore-store TTL. Digest originals in ~/.distil/restore/ now age out after DISTIL_RESTORE_TTL_DAYS (default 14, 0 disables), on top of the existing 500-file count cap — plaintext agent content no longer sits on disk indefinitely under low traffic.
  • Windows CI job (windows-latest, 3.12) — the classifier says OS Independent; now the fcntl/termios/SIGPIPE guards are actually exercised.

Fixed

  • Gateway accounting is crash-safe. Per-tenant counters now checkpoint to disk at most every 30 s during traffic (atomic replace), not only in the graceful-shutdown path — a kill -9/OOM loses ≤ 30 s of accounting instead of everything since startup.
  • MCP store race. Two concurrent distil_compress tool calls could interleave load/save and silently drop one handle (later distil_expand on it failed). The read-modify-write now runs under an advisory lock, same pattern as the savings ledger.
  • distil wrap proxy self-heals. If the embedded proxy's accept loop ever crashes, it logs and restarts instead of leaving the wrapped agent with connection-refused for the rest of the session.

Changed

  • Claims tightened to what the artifacts back (audit follow-up): "only Distil certifies the reversible tier" (Headroom ships an uncertified retrieve — the old wording overclaimed "offers"); the ~1,000× speed multiple now carries its no-ML-model-vs-transformer-inference framing at first mention; LLMLingua-2's SWE-bench row notes only 16/500 runs completed; the Rust core is labeled build-from-source (published wheels run the pure-Python engine).

1.10.1 — 2026-07-05 — Review follow-ups

Fixed

  • No more 0-savings ledger rows. A flush window that saved nothing (typical under --lossless-only) no longer writes a ledger row — session ledgers stay signal-only. Request accounting elsewhere is unchanged.
  • Dedup markers are expand-recoverable. The «repeat of earlier tool output …» marker now carries the handle= form that distil expand keys on, resolving to the byte-exact original.

Changed

  • One statusline label for decision-equivalence. de 12/25 while collecting → ✓de 99.5% (n) at 25+ samples (was eq for the rate form).

1.10.0 — 2026-07-05 — Production hardening: truly lossless, honest accounting, lifecycle fixes

Fixed

  • --lossless-only is now truly lossless (no Tier-1 stubs). Previously a Tier-1 reversible digest stub could appear in a lossless-only session — but without an injected expand tool the agent could never recover it, making the stub effectively irreversible. The flag now folds directly into verbatim (Tier-0-only) at all three proxy entry points (aproxy, proxy, gateway). No separate --verbatim flag needed.
  • Recoverable digests everywhere. All four digest forms (Tier-1, columnar, template, skeleton) now emit handle= markers backed by the RestoreStore. Originals persist to ~/.distil/restore/ (respecting DISTIL_HOME) and survive proxy restarts — distil expand <handle> works across sessions.
  • Honest savings accounting (numbers dip — that means they became correct). Records are booked only after a confirmed 2xx response; failed or retried requests are no longer counted. New ledger rows carry acct:2; mixed-era ledgers print a footnote: "(includes N records from pre-1.10 accounting — savings may be overstated)". The cache simulator is now write-once-then-read, eliminating a double-counting path.
  • Terminal corruption fixed. distil wrap now saves and restores the terminal state (termios) on exit, so a wrapped agent that dies mid-output no longer leaves the terminal in raw mode.
  • Upgrade version-skew warning. distil upgrade now detects running proxy/wrap/gateway processes and warns to restart them — a live proxy loaded pre-upgrade modules can hit a version-skew crash on lazy import mid-request.
  • Decision-equivalence in statusline. The ✓/⚠/✗ eq <rate>% (n) display now appears only at ≥ 25 shadow samples; below that threshold it is suppressed everywhere (statusline, leaderboard, doctor, dashboard) — a rate over a handful of samples is noise, not a guarantee.
  • Ledger resilience. Corrupt lines are tolerated (skipped with a warning rather than crashing the whole read), cross-process writes use advisory file locking, and a backup is kept on each write cycle.
  • Gateway persistence and tenant cap enforced at state load. Per-tenant accounting persists across restarts; the tenant cap is checked at state-load time, not only at request time.

Added

  • tests/test_live_certified_equivalence.py — pins the live proxy's compression decisions to the certified strategy, making any drift a visible, reviewed change. The one documented intentional delta: a recency carve-out keeps the last few tool-result turns verbatim so the agent always sees its freshest output byte-exact.

1.8.1 — 2026-07-04 — Believe-it UX + honest ▼0

Fixed

  • Statusline session view: shows THIS session first (▼75.0K −62% $0.31), lifetime as one Σ figure; theme-proof 256-color palette + ✓/⚠/✗ health glyphs (basic magenta rendered unreadable on dark themes); a session with traffic but nothing trimmed yet reads watching · N seen, not ▼0 −0%.
  • ▼0 self-explains: every compressed response carries x-distil-mode + x-distil-compressible-tokens; distil doctor warns when the always-on proxy runs in verbatim (which caps savings near zero by design).
  • Ledger records carry a session id; proxy no longer writes zero-baseline records.

Added

  • Landing hero: two-door router + a real terminal proof card; benchmark chart in the hero; site-wide editorial layering (both audiences, less prose).
  • distil stats --badge (shareable measured-savings badge); LAUNCH.md.

Docs / process

  • E14 propagated to ALL paper artifacts (main.pdf, NeurIPS variant, PAPER.md); paper-build now rebuilds + commits PDFs on push to main so they can't drift.

1.8.2 — 2026-07-04 — GA polish: no papercuts

Fixed

  • No raw tracebacks on bad input. A missing/malformed input file across 8 commands dumped a Python stack trace; one guard at the dispatch point now prints a clean distil <cmd>: <error> and exits 2.
  • --help no longer lists commands that don't exist (expand/sweep/gate/ corpus/adaptive were phantom); a regression test fails if any return.
  • One installer-detection source of truth (onboard.install_method): upgrade, offboard, and doctor all use it, so upgrade/uninstall hints are always the runnable command (brew/pipx/uv/pip-with-venv-caveat) — no more bare pip that PEP 668 blocks.
  • distil doctor detects shadowed installs (two distil on PATH) and verbatim mode — the two traps that made "▼0" or "upgrade didn't take".

Added

  • distil version (the word people type) and distil upgrade (installer-aware).
  • World-class README hero (runnable terminal proof block); figures in PAPER.md; Homebrew tap auto-bumps on release; GHCR image + PDFs auto-rebuilt on push.

1.8.3 — 2026-07-04 — Latest & greatest: statusline, plain-English docs, self-service

Added

  • Redesigned status line — rich by default (distil · session ▼7.8K · 4% smaller · $0.31 · total ▼27.0M · ✓eq 99%), the session number pops in bold green; DISTIL_STATUSLINE=minimal for crowded composite lines. Clear session/total labels, N% smaller (no misleading −), cohesive teal/green palette.
  • distil version and distil upgrade (auto-detects brew/pipx/uv/pip).
  • Landing page: a plain-English "How it works" section for non-technical readers.

Fixed

  • distil doctor flags shadowed installs (two distil on PATH) and verbatim mode.
  • distil offboard prints the uninstall command that actually works per installer (no bare pip that PEP 668 blocks).
  • No raw tracebacks on bad input; --help no longer lists commands that don't exist.
  • One installer-detection source of truth (onboard.install_method).

Docs

  • Lean README (~40% less prose) + live/clickable badges; 18-page site polish; every link verified; PAPER.md figures; honest banner.

1.8.4 — 2026-07-04 — Statusline polish + landing/docs GA audit

Changed

  • Status line fully colored (cohesive teal→green, no gray): session number pops bold green, trim rate mid-teal, total muted teal. N% smaller (not a misleading −N%).
  • Version single-sourced — __init__ reads pyproject instead of a hardcoded literal (no drift, no merge-back conflicts).

Fixed (docs, proactive audit)

  • Landing page: Python 3.11+→3.9+ (factual); heading hierarchy; two "How it works"→ one is "Under the hood"; proof section now cites E14 (42.0% vs 39.2%); plain-English section linked from nav + hero; smart quotes.
  • benchmark.html cites E14; getting-started smart quotes + stale version example.

1.8.5 — 2026-07-04 — Statusline clarity + self-diagnosing doctor

Fixed

  • Statusline no longer flickers across terminals. distil default spawns a proxy+session per terminal; the live ▼ now aggregates a 15-minute activity WINDOW across ALL sessions instead of one flickering "latest session".
  • Zero-savings state is unmistakable: ✓ on · waiting for a large read (bright green, clearly active) instead of a dim, easily-misread "watching".
  • distil doctor self-diagnoses the two traps: live routing warns when a wrap/proxy is running but no traffic is recorded (agent bypassing distil); this session explains the watching state.

1.8.6 — 2026-07-04 — GA presentation + full-surface audit

Rendered every user-facing surface and fixed everything found — the engine was already proven solid (an evidence-based runtime audit came back clean).

Fixed — presentation & consistency

  • Status line: ONE pattern in every state (distil · <live> · total ▼<lifetime>); live = 15-min window across ALL terminals (no session flicker); zero-savings reads ✓ on · waiting for a large read (never a broken-looking ▼0 −0%); all-teal palette, no gray.
  • No tracebacks on bad input anywhere — added NotADirectoryError to the dispatch guard (a --corpus pointing at a file leaked a traceback on 6 commands); perf --iterations 0 and holdout --control-fraction out of range now give clean errors; ingest no longer silently 'succeeds' on garbage.
  • decision-equivalence suppressed below 25 samples EVERYWHERE (status line, leaderboard, doctor, dashboard, shadow-stats) — no 100% guarantee off n=1.
  • Dollars 2dp (or notional on a subscription); correct singular/plural (1 request/1 sample/1 matched trajectory); online shows 87.3% not 16 digits; certify p=<0.0001 not p=0.
  • distil default now says: RESTART your agent — the #1 onboarding trap (an agent started before the alias bypasses distil → savings stay at zero). distil doctor also flags this (live routing) and explains the watching state (this session).

Docs

  • Statusline state table (saving / watching / idle) in README + Integrations; proof-first hero everywhere (dropped the unmeasured "in half"); technique numbering aligned CLI↔site.

1.9.1 — 2026-07-04 — Quiet client disconnects

Fixed

  • No more traceback spam on client disconnects: agents (Claude Code especially) reset/abandon connections constantly — cancelled streams, retries, statusline polls — and every one dumped a full ConnectionResetError: [Errno 54] stack trace into the terminal running distil wrap/proxy/gateway. All three servers now run on a QuietHTTPServer that silently drops ConnectionResetError / BrokenPipeError / ConnectionAbortedError; real errors still print.

1.9.0 — 2026-07-04 — Per-session savings + hardened CI

Added

  • True per-session status line (the headline UX): each terminal now shows ITS OWN session's savings (distil · ▼30.0K · 60% smaller), while total stays lifetime across all sessions. distil wrap stamps a DISTIL_SESSION id inherited by both the proxy (which tags every ledger record) and the agent → the status line it spawns — so attribution is exact, with no cross-terminal bleed. A fresh terminal reads ✓ on until it compresses something. The distil dashboard mirrors the same per-session view.

Changed

  • Sharper positioning everywhere (README, docs site, social image): dropped the "statistical fidelity certificate" jargon → "Every other compressor asks you to trust it won't break your agent. Distil is the only one that proves it won't." The E14 result reframed as a win — "compressed context didn't just match the full context — it beat it: 42.0% vs 39.2%."

Quality

  • 95% test coverage (was a 78% floor), 1140+ tests: the CLI, status line, doctor, ledger, the network layer (proxy / gateway / streamrelay / async proxy), and the statistical-certificate paths are all exercised. Genuinely external code (the torch training loop, live-model proof-harness runners) is documented-and-omitted, not hidden.
  • Fixed a Python-3.9-only flaky proxy-timeout test and a $HOME-dependent test that failed in a clean CI environment; the coverage floor now ratchets at 95%.

1.8.0 — 2026-07-04 — GA: compression that beats full context, certified

Headline result (E14, SWE-bench Verified n=500, official harness)

  • The v1.7 surprise-preserving digest resolves 42.0% of tasks vs full context's 39.2% (+2.8pp, paired CI [−0.6, +6.2]pp — statistically non-inferior with the point estimate above full) and +5.2pp over the E8 head-digest gate. The shipped trajectory certificate certifies it (α=0.10, observed degradation 6.2%). Mechanism confirmed end-to-end: keeping a traceback's tail preserves the anomaly the next action needs. Paper §E14; docs/compare.html.

Added

  • GA container image: ghcr.io/dshakes/distil (amd64+arm64), published on release tags. Multi-stage, non-root, gate-verified.
  • Session-first statusline: this session leads (▼75.0K −62% $0.31), lifetime collapses to Σ27.0M; compact composite-friendly grammar; theme-proof 256-color palette with ✓/⚠/✗ health glyphs; equivalence shown only at 25+ shadow samples (a rate over a handful of samples is noise).
  • distil stats --badge — shareable shields.io badge of your measured savings; ledger records carry a session id (ledger.summary(session=)).
  • Decision-equivalence + session cards on the HTML savings page and a session row in the TUI dashboard.
  • Claude Code plugin 1.8: /distil-certify and /distil-badge commands; full command table on the Integrations page.
  • E14 benchmark condition (distil_gated_surprise) + committed results, paper section, and macro generator.
  • docs/compare.html (honest head-to-head), LiteLLM Proxy recipe, compliance-teams section, THREAT_MODEL.md, LAUNCH.md.

Fixed

  • Homebrew tap served 0.24.0 (pre-GA) — bumped to current and verified.
  • ledger.default_path() honors DISTIL_HOME; forward path never follows redirects; identity encoding on compressible requests; 411 on chunked bodies; gateway stops echoing the anon tenant hash; mypy-clean package with typecheck + coverage floor in CI.

1.7.0 — 2026-07-03 — The trajectory-level certificate, true streaming, trust-critical savings fixes

Added

  • Trajectory-level risk certificate (distil certify-trajectories, distil.certify.trajectory_risk): certify the invariant that actually transfers to task success — a distribution-free Conformal Risk Control / Learn-Then-Test bound on end-to-end task degradation over matched full-context/compressed runs, with stated exchangeability assumptions, small-sample refusal, and an anytime-valid drift monitor that flags when the certificate needs recalibration. This is the corrected certificate target: per-step next-action equivalence provably overpredicts multi-step success (our E7 experiment; arXiv 2412.17483).
  • Outcome-guided compression policy (distil.compress.guideline): ACON-style learning from trajectory outcomes — content classes whose digestion co-occurs with end-to-end regressions get protected byte-exact. Never-regressing by construction (only makes compression more conservative); content-free signatures only; always on in the proxy.
  • Surprise-preserving retention: a fourth salience signal — error lines, failures, anomalies, and unified-diff changes are over-retained (the "lost if surprise" failure mode of lossy compressors), plus file-path protection.
  • True streaming pass-through in all three servers (proxy, async proxy, gateway): SSE responses relay chunk-by-chunk, preserving time-to-first-token (previously every response was buffered start-to-finish). Shadow-mode decision-equivalence accounting tees off the streamed bytes.
  • --json output on doctor, leaderboard/stats, and shadow-stats; a stats alias for leaderboard; grouped distil --help.
  • Doctor checks for pricing-catalog drift (unpriced models in the ledger) and tokenizer grade (heuristic vs billing-grade counts).

Fixed

  • Savings were priced at one fixed model. The proxy now accounts each request under the model it names (mixed Opus/Haiku sessions are no longer all priced at the Opus rate), the pricing catalog covers current model ids (dated/Bedrock/Vertex shapes resolve too), and unknown upstreams (e.g. Gemini) record token savings with dollars=0 rather than being silently billed at Claude rates. The async proxy now records savings at all.
  • One-liner def f(): pass functions vanished from code skeletons, leaving orphaned ... where the signature should be.
  • --shape-output broke against the Anthropic API (injected role:"system" into messages, which /v1/messages rejects); the directive now goes into the top-level system field on Anthropic bodies.
  • Upstream calls had no timeout — a wedged upstream pinned a worker thread forever; now a finite (env-tunable) timeout maps to a 504.
  • Savings flushed only every 50 requests and were dropped on kill — now every 10 requests or 30 s, plus a SIGTERM handler that flushes (and forwards the signal to the wrapped agent).
  • Gateway tenant identity trusted a client header — accounting identity now derives from the credential hash; x-distil-tenant is honored only under --trust-tenant-header. /distil/stats and /distil/dashboard require --admin-token (Bearer) and are refused on non-loopback binds without one. The MCP handle store is bounded and chmod 0600.
  • Concurrency race in expand-mode learning stats (intermittent 500s), sparse record arrays no longer fold ambiguously, delta replay order is a declared field with a loud error on mismatched turns, salience re-injection keeps indentation, online warns when reporting train-set metrics.

Changed

  • Status line is now glanceable: shows the percent trimmed next to the token figure (1.2M→0.5M tok −58%), a single $X.XX saved delta instead of two dollar figures, and colors decision-equivalence by health (green ≥99%, yellow ≥95%, red below) so a fidelity regression is visible at a glance.
  • distil stats now prints the orig→compressed token totals with the percent trimmed and the live decision-equivalence (with shadow sample count) alongside the dollar totals.

1.6.2 — 2026-06-30 — Consistent version reporting

Fixed

  • distil --version (and distil doctor) reported 1.6.0 on the 1.6.1 release. The version lived in two places and only pyproject was bumped, so the published wheel's __version__ lagged. Now single-sourced: distil.__version__ reads the installed distribution metadata (importlib.metadata), so the CLI can never drift from the published package again. 1.6.2 carries the same Python 3.9+ fix as 1.6.1.

1.6.1 — 2026-06-30 — Installable on Python 3.9+ (fixes "from versions: none")

Fixed

  • pipx install distil-llm / pip install distil-llm failed with Could not finda version that satisfies the requirement distil-llm (from versions: none) on stock macOS. Root cause: requires-python was >=3.11, but macOS ships Python 3.9 as the system python3, so pip filtered out every release and reported that misleading message. The package is stdlib-only and uses from __future__ import annotations, so it already imports and passes the distil bench gate on 3.9/3.10 — the floor was simply set too high.
  • Lowered the floor to Python 3.9 (requires-python = ">=3.9", classifiers added) and aligned distil doctor's version check. CI now runs the full suite + gate on 3.9–3.13 so the support claim stays true. Docs/troubleshooting updated. > Reaches users once 1.6.1 is published to PyPI (the live 1.6.0 still pins > >=3.11). Publish by pushing a v1.6.1 tag.

1.6.0 — 2026-06-30 — Onboard ensures everything

Added

  • distil onboard now ensures you have everything — including a permanent install. When run ephemerally (e.g. uvx --from distil-llm distil onboard), distil isn't on PATH, so onboard detects that and offers to install distil permanently first (pipx/uv/brew, per your machine) before wiring the status line and routing your agent. Makes uvx --from distil-llm distil onboard a true one-command setup. Intelligent by default — no flag to opt in.

1.5.0 — 2026-06-30 — Clean teardown

Added

  • distil offboard — remove distil's footprint, the inverse of onboard. Undoes the shell default (alias/env block), stops + removes the always-on proxy service, and unwires the status line from Claude Code settings — asking before each (non-interactive without --yes removes nothing). Your savings ledger is kept unless you pass --purge. It can't uninstall the running package itself, so it prints the exact uninstall command for how distil was installed (pipx/uv/pip). distil default --undo now also stops a running proxy service (launchctl/systemctl), not just deletes its definition file.

1.4.0 — 2026-06-30 — Make distil the default

Added

  • distil default — make distil the default for your agent, no per-session distil wrap. Writes a single managed (marked, backed-up, idempotent) block to the shell rc that distil actually detects for this machine — zsh (.zshrc), bash (.bashrc/.bash_profile), fish (config.fish), or PowerShell ($PROFILE) — using the right syntax for each (alias / function / export / set -gx). An explicit $SHELL wins over file-existence guesses, and the command reports what it detected rather than acting blind. --always-on installs a persistent proxy service (launchd / systemd) + ANTHROPIC_BASE_URL so every SDK routes through distil (with an honest single-point-of-failure caveat); --undo removes whichever is installed. distil onboard now offers it interactively.
  • distil onboard is now upgrade-aware and agent-ready. It checks PyPI (offline-safe) and, if a newer release exists, shows the exact upgrade command for your install method (pipx/uv/pip) — --upgrade runs it. New distil onboard --json emits the full environment + version status + recommendations as structured data so an agent can reason over it.
  • Intelligent /distil-onboard skill — rather than a static installer, the Claude Code command now senses via --json, assesses your situation (upgrade, which agent, billing reality, gaps), and guides you through setup + validation conversationally, asking and adapting rather than ticking boxes.

1.3.0 — 2026-06-30 — One-command onboarding

Added

  • distil onboard — one command that detects your environment (OS, package managers, agent CLIs, install method, the anthropic extra, Claude Code + subscription), wires the savings status line, and prints a next-steps guide tailored to what it found — how to route the detected agent (subscription-safe vs metered), validate outcomes with shadow mode, watch savings, run the gate, and re-verify with distil doctor. --dry-run changes nothing. Cross-platform (macOS / Windows).

1.2.0 — 2026-06-30 — Setup & diagnostics UX

Friction-killers for getting distil running and trusting it.

Added

  • distil doctor — one command diagnoses a setup end-to-end: distil/Python version, savings ledger (subscription-aware), shadow-validation status, an in-process proxy round-trip self-test (proves the proxy machinery works with no network), the optional anthropic extra + API key, and Claude Code status-line wiring + subscription detection. Exits non-zero on any failure.
  • distil setup — wire the savings status line into Claude Code's settings.json in one command: idempotent, never clobbers an existing line without --force (backs it up), preserves all other settings.
  • Subscription auto-detect — the status line and dashboard now drop the notional dollar figure automatically on a Claude OAuth subscription (no more manual DISTIL_SUBSCRIPTION=1; the env var still overrides).
  • Status line shows the shadow sample count next to eq% (eq 99.5% (1.2k)) so the confidence is visible.
  • Dashboard gains a live recent-decisions strip under decision-equivalence (▰ same next action · ▱ changed), refreshing with the panel.
  • Verified multi-provider shadow — decision discrimination tested for Anthropic / OpenAI / Gemini response shapes.

1.1.0 — 2026-06-30 — Hardening + live-validation UX

Post-GA hardening of the 1.0 line, validated end-to-end across every command. Zero-dependency stdlib core; 665 tests.

Fixed

  • Status line BrokenPipeError — on Python 3.13+, when the status-line consumer (e.g. Claude Code) read the line and closed the pipe, the interpreter's shutdown flush faulted with a traceback. The statusline path now flushes under guard and exits cleanly; verified 0/40 on the real binary.
  • Shadow-mode dropped samples — each sampled decision ran in a daemon thread that was killed on proxy teardown (quick runs / last turn), so distil wrap --shadow could show 0 samples despite live traffic. In-flight comparison threads are now drained (bounded) on shutdown.
  • Raw tracebacks → actionable messages — --tokenizer/--runner anthropic (missing anthropic extra or API key) and distil ingest --input <bad-path> now fail with a clear, single-line message instead of a Python traceback.
  • Claude Code plugin manifest — repository must be a string URL (was an object), which blocked installation.

Added

  • distil dashboard — a live, zero-dependency terminal TUI: alternate-screen framed panel with Unicode bars for token-trim and decision-equivalence, original → compressed tokens/cost, and per-trajectory bars.
  • distil wrap --shadow RATE — one-command live decision-equivalence: wraps the agent, starts the proxy, sets the base URL, and shadow-samples — no second terminal, no manual env var.
  • Status line now shows original → compressed tokens and cost, surfaces live decision-equivalence (eq N%) when shadow has samples, and drops the notional dollar figure on flat-rate subscriptions (DISTIL_SUBSCRIPTION=1).
  • Plugin commands — /distil-stats, /distil-shadow, /distil-dashboard alongside /distil.
  • Docs — README and the docs site document --shadow outcome validation, the dashboard, subscription mode, and the one-command shadow flow.

1.0.0 — 2026-06-29 — General Availability

1.0 / GA. The compression engine, the proxy/SDK integrations, and the decision-equivalence certificate machinery are production-grade, API-stable, and covered by 658 tests with a zero-dependency stdlib core. This release folds in the cross-model, cost-frontier, and continuous-assurance work that landed after 0.28.0 and declares a stable public surface.

What "1.0 / GA" means (and what it doesn't). It is a commitment to a stable API and to the contract that protects you — certify decision-equivalence, or fall back to full context; never silently lossy. It is not a claim that aggressive compression is safe on every agent untuned: E7/E11 show the opposite, which is precisely why the operating point is auto-calibrated per deployment and fail-safe. Honest scope, unchanged: the guarantee is distribution-free and finite-sample, conditional on exchangeability with your calibration distribution. See docs/GA_READINESS.md for the full ledger of what is closed and what remains empirical breadth.

Added — cross-model generality (E11)

  • Validated across 5 models / 3 vendors. The long-horizon harness (30-turn ReAct, SWE-bench Verified) now reports gpt-4o-mini and gpt-4.1 (OpenAI), Sonnet 4.6 (Anthropic), Haiku 4.5 (Anthropic, n=500), and DeepSeek-V3 (n=200). gate@12 shows no statistically significant degradation on any of the five models. The two well-powered runs (Haiku n=500, DeepSeek n=200) confirm non-inferiority; the three n=50 runs are directionally consistent with wide CIs (honestly marked as not powered).
  • Corrected finding. An earlier reading of DeepSeek alone ("aggressiveness must scale with model capability") is refuted by the wider sweep: harm appears only as the product of realized compression × the agent's reliance on aged-out context — a workload×model interaction, not raw capability. A fixed gate_recent cannot predict it, which is why you must calibrate on outcomes per deployment.
  • OpenAI 429 handling — retry on TPM rate-limits with backoff + Retry-After.

Added — auto-calibration, productionized (closes the headline GA risk)

  • distil calibrate selects the most aggressive working-set size whose task-success loss is non-inferior to full context (paired McNemar), and fails safe to full context if none certifies — the operating-point analogue of the certificate. Reproduces the manual E11 choice automatically (selects gate@12, rejects gate@6 on DeepSeek). distil/calibrate.py, tests/test_calibrate.py. The relevance gate is now a shippable library primitive (distil/gate.py: working_set_indices, gate_fraction), not benchmark-only.

Added — cost frontier under the motto (E12)

  • Cache-monotone gate (gate.py:monotone_gate) — deterministic append-only digests so the digested prefix is byte-stable and prompt-cache/KV reuse captures it.
  • Graded gate (gate.py:graded_gate) — per-distance compression tiers, certified with the tighter empirical-Bernstein (Maurer–Pontil) bound (conformal.py).
  • Speculative expansion (speculative.py) and constrained-bandit operating-point search (calibrate.py:bandit_select_operating_point) — fail-safe, shipped + tested. All levers cut cost inside the certified envelope; they never trade the guarantee for dollars.

Added — continuous assurance under drift (E13)

  • Anytime-valid drift monitor (drift.py:DriftMonitor) — a betting e-process for H0: risk ≤ α (Waudby-Smith & Ramdas 2023) you may check after every turn with false-alarm probability ≤ δ regardless of how often you peek (Ville's inequality). Trips when live decision-change exceeds the certified budget → recalibrate or fall back.
  • Cross-family grader ensemble (ensemble.py:EnsembleGrader) — conservative "any-change" aggregation keeps measured risk an upper bound even if one grader family is unfaithful.
  • Anytime-valid certificate for graded losses (conformal.py:betting_upper_bound).

Changed

  • Package version reconciled to 1.0.0 (pyproject.toml, distil/__init__.py, CITATION.cff); PyPI classifier → Production/Stable.
  • Docs/site test counts corrected to 658; the landing page's E11 narrative updated to the corrected (5-model) finding.

0.28.0 — 2026-06-26

E10: trajectory-level decision-equivalence certificate — the first distribution-free, out-of-sample-proven guarantee at the whole-run level for agent context compression.

  • E10 trajectory-level certificate. Lifts the per-turn E2 certificate to the full trajectory (task) level using the same Learn-Then-Test / Hoeffding–Bentkus engine (distil.conformal.certified_risk_bound), inverted to a (1−δ) upper confidence bound on per-trajectory 0/1 loss. Two loss functions on the full 500-instance SWE-bench Verified set (δ=0.05):
    • Divergence (outcome ≠ full context): empirical 14.4%, certified ≤ 18.0%.
    • Harm (full resolved the task, gated did not): empirical 8.4%, certified ≤ 11.4% — about 1 in 9 solvable tasks, certified.
    • Plain-language: "With 95% confidence, the relevance-gated compressor changes a run's outcome on ≤18.0% of exchangeable tasks and costs a solvable task on ≤11.4%."
  • Out-of-sample proof. Over 1000 random calibration/test splits, the bound β is certified on the calibration half and checked on the disjoint test half. Realized coverage: 95.4% (divergence) and 96.7% (harm) — both at or above the 95% target. The bound holds on held-out data, not merely asserted on training data.
  • Honest reporting: ungated reversible tier. The ungated tier (condition D, E8) also certifies: divergence ≤23.2%, out-of-sample coverage 93.9% — marginally below the 95% target. Reported without softening.
  • Honest scope. The guarantee is exchangeability-conditional: valid for traffic exchangeable with the calibration distribution (SWE-bench Verified, this agent + model). Changing the agent, model, or task distribution requires re-certification.
  • Why it matters. E2 guaranteed a per-turn proxy. E7/E8 showed that proxy doesn't naively transfer to task success under aggressive compression. E9 quantified the composition gap. E10 closes it: the first trajectory-level, distribution-free decision-equivalence certificate for agent context compression (to our knowledge).
  • Reproducible. benchmarks/trajectory_certificate.py; numbers trace to docs/paper/results/swe_e2e_longhorizon/trajectory_certificate.json.
  • Docs updated: docs/research.html (E10 section with results table and OOS proof), docs/index.html (honest-scope headline line), docs/concepts.html (certificate callout).

0.27.0 — 2026-06-26

Final E8 long-horizon results: 6-condition frontier including Headroom competitor, skeleton digest, sticky expansion, digest-mode-per-tier ablation, and the E9 trajectory-composition certificate bound.

  • E8 long-horizon SWE-bench Verified — final 6-condition frontier. A custom multi-turn ReAct coding agent (read / search / edit_file / run_tests, up to 30 turns, claude-haiku-4-5, temp 0) run end-to-end on the full 500-instance SWE-bench Verified set, scored by the official swebench harness (hidden tests, per-instance Docker). Runs average ~27 turns. Six conditions, same agent, compressor differs (ordered by pass@1, Wilson 95% CI, resolved/500):
    • A (full context): 196/500 — 39.2% [35.0, 43.5]
    • E (distil reversible, relevance-gated): 184/500 — 36.8% [32.7, 41.1]
    • F (Headroom, lossy competitor): 163/500 — 32.6% [28.6, 36.8]
    • D (distil reversible + skeleton digest, ungated): 162/500 — 32.4% [28.4, 36.6]
    • B (distil trunc@500, aggressive lossy): 28/500 — 5.6% [3.9, 8.0]
    • C (LLMLingua-2, lossy competitor): 12/500 — 2.4% [1.4, 4.2]
    • Total API spend across all six conditions: $571.15
  • Key results (paired McNemar, same 500 instances).
    • Gate (E) vs full context (A): −2.4 pp, 95% CI [−5.7, +0.9], McNemar p=0.19. Non-inferior at a 6 pp margin (borderline at strict 5 pp). This is a non-inferiority result, not equivalence. The gate is the only condition statistically non-inferior to full context.
    • Gate (E) vs Headroom (F): +4.2 pp, McNemar p=0.035. Statistically significant. Distil is not cheapest — Headroom is cheaper — but beats Headroom on task success with significance.
    • Gate (E) vs LLMLingua-2 (C): 174 gate wins vs 2 LLMLingua-2 wins, McNemar p<0.001. E and C remove nearly identical context fractions (53% vs 52%), isolating what is kept as the deciding factor.
    • Lossy truncation (B) vs full: p<0.001.
  • Honest headline. On the axis that defines the field — certified decision-equivalence plus real task success — distil leads. It does not claim cost-domination. Headroom is cheaper. The claim is: the only certified and reversible compressor, with the highest task-success of any compressor tested, and the only one statistically non-inferior to full context.
  • New technique: content-aware skeleton digest (distil/compress/skeleton.py). For the active-recovery (ungated) tier, large source files are digested to a navigable skeleton: every import/class/def signature retained, traceback tails kept, bodies elided. Deterministic and stdlib-only (no model, no network — auditable and secure). Byte-exact reversible via content handle. Lifted ungated pass@1 from 28.8% to 32.4% (condition D).
  • New technique: sticky expansion (distil/expand.py). Once the agent recovers a block via distil_expand, that block stays full for the rest of the session (handles are deterministic). Eliminates re-expansion thrash on repeatedly-accessed files. Never-regressing by construction.
  • Honest ablation: digest mode per tier. Applying the skeleton digest to the relevance-gated (passive) tier regressed pass@1 from 36.8% to 5.6%, matching lossy truncation. A navigable digest makes the agent over-trust the summary and stop re-reading. Skeleton digest is correct for the active-recovery tier; head-truncation is correct for the passive tier. This finding is published as-is.
  • E9 trajectory-composition certificate bound. The per-turn certificate extends to multi-turn trajectories. Across ~27-turn runs, only ~1.8 turns are outcome-determining, so the naive composition bound (which becomes vacuous at ~27 turns) overstates risk. The formal per-trajectory bound remains an open problem; reversibility is the operative safety guarantee for the active-recovery tier.
  • Docs updated: docs/research.html (6-condition table, Headroom row, skeleton/sticky sections, honest-ablation note, certificate scope), plus docs/index.html, docs/concepts.html, docs/benchmark.html, docs/techniques.html (skeleton digest and sticky expansion sections).
  • Numbers trace to docs/paper/results/swe_e2e_longhorizon/swe_bench_verified_longhorizon.json.

0.25.1 — 2026-06-25

Version bump only; same content as the v0.25.0 release notes — fixes the PyPI publish that failed on a duplicate filename in v0.25.0 (the package version was still 0.24.0, so the wheel/sdist collided with an already-uploaded distribution). The v0.25.0 tag and GitHub Release are intentionally left in place. See the v0.25.0 release for the substantive change (Phase 5 / E7 SWE-bench Verified end-to-end eval).

0.24.0 — Ecosystem hooks + on-motto gap-closing

New surface area for agent frameworks and observability — every addition kept under the decision-equivalence certificate, with the platform scope-creep deliberately declined.

  • LangGraph hook (distil/integrations/langgraph.py) — a drop-in pre_model_hook() that compresses graph state right before the model node, plus a compress_state() helper for manual use inside any node. Duck-typed (never imports langgraph/langchain); returns only the updated message list so every other state field is untouched. Joins the existing LiteLLM + LangChain hooks. Example: examples/python_langgraph.py.
  • Cache-prefix observability — the proxy now emits x-distil-cache-prefix-msgs: <n> under --session-delta, exposing exactly how many leading messages stayed byte-identical vs the previous turn (the prompt-cache-read region). The verifiable benefit of a prefix-freeze router, content-free — distil is cache-monotonic by construction, so the prefix is real, not rewritten.
  • Pluggable salience scorer seam — salient_tokens(..., scorer=…) accepts an optional callable (a semantic / NER / embedding model) whose spans are unioned into the model-free signals. Off by default (runtime stays model-free, zero-dep); a bad scorer can never break compression (guarded), and whatever it returns is still judged by the same certificate — the seam adds coverage, never an unverified guarantee.
  • Docs: README now documents the framework hooks and a "Deliberately not a platform" section — why memory/knowledge-graph, hosted semantic cache, and editor-auth are out of scope (they can't be put under the certificate), and what we adopted instead because it survives the gate.

0.23.2 — Mobile docs, animated architecture diagram, distribution fix

  • Fixed a broken Homebrew distribution. Both formulas (repo + tap) had frozen their url/version at v0.21.0 while the sha256 advanced — so brew install failed on a sha mismatch. Root cause (a version-specific regex in the update step) fixed with a version-agnostic pattern; both formulas now consistent at the current release and verified against the published tarball.
  • Mobile-responsive docs. Wide benchmark tables now scroll horizontally instead of overflowing; landing stats collapse to one column, CTAs stack, padding/typography scale down — across the docs site and the landing page.
  • New animated architecture diagram (docs/assets/architecture.svg) — a realistic depiction of the pipeline (agent → compress/cache-pin/forward → provider), the transparent recovery loop, and the quality-contract band (certificate · shadow · flywheel), with flowing-data animation. Shown on README, Concepts, Architecture.
  • Vocabulary consistency: distil bench now reports savings as "reversibly" (the strategy uses the Tier-1 reversible digest), matching the v0.23.1 terminology.

0.23.1 — Honest vocabulary: "reversible" vs "lossless"

A precision pass on terminology so no claim can be read as an overclaim:

  • The default Tier-1 digest is now described as "reversible" (byte-recoverable on demand), not "lossless". "Lossless" is reserved for the byte-in-context tier (Tier-0 / --verbatim), where the model sees content unchanged. "Lossy" stays for the irrecoverable competitors. All three Distil tiers remain certified decision-equivalent. Updated the README headline + prose, the benchmark method label, and added an explicit three-tier definition to the README and Concepts page.
  • --safe added as a clearer alias for --lossless-only (the policy/subscription-safe mode: no lossy shaping, no tool injection — the reversible digest still runs); --verbatim remains the byte-in-context switch. Internal strategy/ladder identifiers are unchanged (no behavior change).

0.23.0 — GA polish: grounded docs, genuine head-to-head, recipes

A go-live pass: every customer-facing claim audited against code, every benchmark number re-measured genuinely, and the docs made white-glove.

  • Genuine, apples-to-apples benchmarks. Fixed the local competitor adapters so they actually engage: Headroom is now driven its real whole-conversation way (optimize=True) instead of no-op'ing on a per-block user message; LLMLingua-2 is applied to every tool result and memoised (pure function). Corrected the v0.22.0 coding-agent competitor numbers (which were harness artifacts) to the real ones — LLMLingua-2 56.8% tok / 57.2% $ / 274 ms / lossy, Headroom 22.4% tok / −16.8% $ (busts the prompt cache by default) / 5.3 ms / lossy. distil leads on cache-aware dollars (91.1%) and is the only reversible method.
  • Docs claims audit (no fake, all real grounded). Removed fabricated claims (a non-existent "proxy detects subscription keys" path; ingest --format auto-detect; output-savings --mode/--runner flags; invented corpus.validate() invariants), corrected CLI flag tables, fixed compress_messages(verbatim=) signature, the salience module path, and the 8 proxy response headers; clarified that the live 83.2%/53.1%/35.3% run needs an API key and is not offline-reproducible.
  • Diagrams (YC-style, on-brand): cache-delta.svg and ast-delta.svg, embedded in the techniques and benchmark pages.
  • White-glove "Use it on your workflow" recipes in the README — coding and non-coding use cases, a see-it/prove-it table, and a config rule-of-thumb.

0.22.0 — Coding-agent benchmark + two correctness fixes it found

Building the messages-level coding-agent benchmark (benchmarks/codebench.py: read→edit→reread sessions, cache-aware dollars vs the real headroom + llmlingua packages) surfaced two real bugs, both now fixed:

  • Cache-delta is now cache-monotonic by construction. cachedelta.delta_encode was rewritten as a pure per-call walk over the cumulative messages (message i is deduped only against messages 0..i-1). Previously the stable prefix was passed through as originals while the prior turn had emitted those messages as markers, so a re-read flipped marker→original on entering the prefix and busted the prompt cache. The pure-walk encoding emits identical bytes for the cached prefix every turn. (session is now optional — the cumulative conversation is the memory.)
  • Tier-0 never inflates tokens. collapse_runs could turn a run of near-free blank lines into a <<x N>> marker that costs more tokens; the adapter had only a char-based guard. _apply_tier0 now keeps the collapse only when it reduces the token count. (Fixes verbatim mode showing negative savings on whitespace-heavy content.)
  • Net effect on the coding benchmark: verbatim+cache-delta went from −1.4% to a real +43.8% cache-aware savings (reversible); plain verbatim from −3.5% to 0.0%. The PAYG digest remains the dominant lever (~91%). 3 regression tests (498 total).

0.21.0 — Edit-equivalence (decision-equivalence, made precise for code)

  • Edit-equivalence: the decision signature now AST-normalizes code-bearing tool inputs (e.g. an Edit/Write new_str). For coding agents the decision is the edit, so two responses that make the agent write the same code with trivially different whitespace or comments now count as equivalent, while a real logic change still differs. This stops shadow-mode over-reporting drift and lets the certificate claim safe savings it previously, conservatively, could not.
  • Implemented model-free with the stdlib ast (_normalize_decision → ast.dump), applied through shared signature builders so the JSON, streamed (SSE), and chunk-array paths all stay consistent. Non-code strings and non-Python pass through untouched. 5 tests (495 total). ruff clean, verify + bench PASS.

0.20.0 — AST-structural delta (the deepest cache-delta layer)

  • AST-structural delta (astdelta.py, stdlib ast, model-free): for Python, cross-version delta now diffs by parsed structure. Each top-level definition is fingerprinted with ast.dump (attributes off) — invariant to whitespace, comments, and import order. A reformat-only re-read is recognised as "no definition changed" and referenced; only definitions whose AST actually changed are sent in full. Textual diff explodes on reformatting; the structural delta isolates exactly what changed.
  • Wired as the preferred near-duplicate path in cachedelta.py (the --session-delta feature); non-Python or unparseable (mid-edit) source falls back to the textual unified diff, so it never fails a request. Decision-equivalent (unchanged defs are still in cached context) and reversible (distil_expand recovers the full file).
  • 8 tests. Full suite 490 passed, ruff clean, verify + bench PASS.

0.19.0 — Cache-delta context coding (cross-version delta)

The coding-agent moat. The hot path is read → edit → re-read, and the re-read file is a near-duplicate (one hunk changed), so exact-duplicate dedup misses it and re-sends the whole file. Cache-delta coding (cachedelta.py, distil proxy --session-delta, opt-in) sends only the diff:

  • Cross-version delta — a re-read-after-edit is replaced by a reference to the prior version + a unified diff of what changed; exact re-sends become a compact back-reference. Both confined to the volatile suffix — the stable cached prefix is never mutated (cache-monotonicity), so prompt-cache hits survive.
  • Decision-equivalent + reversible: prior-version (still in cached context) + diff carries the same information for the next action; the full current version is kept locally and recovered byte-exact via distil_expand. Shadow mode measures it.
  • Wired into distil proxy / distil wrap (messages format) behind --session-delta; emits x-distil-cache-refs / -delta / -tokens-saved headers. End-to-end a re-read-after-edit saved ~85% of the re-read (902 of 1063 tokens) vs re-sending whole.
  • 10 tests. Full suite 482 passed, ruff clean, verify + bench PASS.

0.18.0 — Streaming-aware shadow mode (Claude Code / Codex / Gemini)

  • Shadow-mode now works on streaming sessions. Real agent sessions (Claude Code, Codex, the Gemini CLI) stream their responses over SSE, which the previous shadow comparison couldn't parse — so it silently recorded nothing. shadow.py now reconstructs the decision from a streamed body: decision_signature_from_body reads a non-streaming JSON body directly and rebuilds a streamed (SSE or chunk-array) one via _decision_from_chunks, accumulating the first tool call across chunks for all three providers (Anthropic input_json_delta, OpenAI tool_calls argument deltas, Gemini functionCall). A streamed response yields the same signature as its non-streamed equivalent, so comparisons are valid.
  • The proxy shadow path now compares raw bodies via decision_signature_from_body, so distil proxy --shadow measures live decision-equivalence on streaming traffic. Verified end-to-end on an SSE tool-call response.

0.17.0 — Decouple compression aggression from auth (--verbatim)

Resolves an overload introduced in 0.16.0. --lossless-only had been redefined to mean "Tier-0 only," which contradicted policy.py (where the reversible digest is the lossless strategy that subscription sessions use) and silently de-tuned autonomous agents on subscription/OAuth from ~70%+ down to ~10%.

  • --lossless-only restored to its policy meaning: lossless strategies only (no lossy output-shaping) + no tool injection. The reversible, certificate-backed Tier-1 digest still runs — consistent with policy.py and the project's definition of "lossless" (reversible + decision-equivalent).
  • New --verbatim flag (proxy / wrap / gateway): skips the Tier-1 digest entirely (Tier-0 only) so the model sees content un-stubbed. The right mode for interactive (human-in-the-loop) sessions or out-of-distribution traffic. Lower savings, byte-in-context fidelity.
  • Adapter/integration kwargs renamed to match: compress_messages(..., verbatim=), compress_generate_request(..., verbatim=); LiteLLM distil_verbatim; LangChain compress_messages(..., verbatim=). Docs reconciled across CLI / adapters / integrations / faq / deploy-security.

0.16.0 — Ecosystem hooks: MCP server + LiteLLM/LangChain

  • MCP server (mcp_server.py, distil mcp): a zero-dependency, stdlib-only Model Context Protocol server over stdio JSON-RPC 2.0. Exposes distil_compress (reversible digest + handle, original kept in a local on-disk store), distil_expand (recover by handle), and distil_savings. Wire it into any MCP client (Claude Desktop, IDEs, agents). The message handler is a pure function and is unit-tested without real stdio; the loop is verified end-to-end.
  • In-process framework hooks (integrations/): LiteLLM (compress/completion/ acompletion) and LangChain (compress_messages, duck-typed over message objects and dicts) compress requests before they leave the process — same reversible compression as the proxy, no sidecar required. Both lazy-import their framework, so distil stays zero-runtime-deps.

0.15.0 — Claude Code plugin + status line

  • distil statusline (new CLI command): renders a compact one-line savings summary from the local ledger (tokens, dollars, runs, and live decision- equivalence when shadow-mode has samples). Reads the optional Claude Code status- line JSON on stdin for the model name; never raises.
  • Claude Code plugin (plugins/distil/ + .claude-plugin/marketplace.json): installable via /plugin marketplace add dshakes/distil. Ships a /distil command (savings report + setup help) and a statusline.sh that calls distil statusline. Honest scope: a plugin cannot reroute a running session or set the main status line from its manifest, so the README documents the one-line settings.json addition; traffic is compressed via distil wrap / distil proxy.

0.14.0 — Google Gemini adapter + true lossless-only

  • Gemini adapter (adapters/gemini.py): the proxy, async proxy, and gateway now compress Google's generateContent request shape (contents / parts / functionResponse) — a third first-class provider alongside Anthropic and the OpenAI-compatible family. text parts get Tier-0 lossless transforms; large functionResponse string values get the Tier-1 reversible digest (recoverable via the local store); functionCall, inlineData, fileData, and model-authored text pass through untouched. Path-detected (:generateContent / :streamGenerateContent), so just --upstream https://generativelanguage.googleapis.com. Shadow-mode live decision-equivalence works for Gemini too. (Expand-tool injection, output shaping, and Gemini context caching remain messages-format-only for now.)
  • --lossless-only is now genuinely lossless-in-context (GA correctness fix). It previously still applied the Tier-1 digest, replacing tool output the model could not recover (tool injection is disallowed on subscription/OAuth) with a stub — despite the "safe for subscription" label. It now applies only Tier-0 transforms in this mode, so the model sees semantically identical content. The aggressive, certificate-backed reversible digest remains the default (PAYG) behavior.

0.13.0 — Shadow-mode live decision-equivalence

  • Shadow mode (shadow.py, distil proxy --shadow RATE, distil shadow-stats): samples a fraction of live requests, runs each one both compressed and uncompressed in a background thread (never blocking the client), and records a content-free live decision-change rate on real traffic. The continuous online counterpart to the offline certificate — decision-equivalence becomes observable in production. Decision = the agent's next tool_use/tool_call; equivalence iff that action matches.
  • README: a "See it working — real-time savings & live equivalence" section (per-request headers, gateway dashboard, genuine-savings ledger, shadow mode, and one-env-var org-wide enforcement).

0.12.1 — GA hardening

Pre-GA security + correctness pass (no behavior change to the happy path):

  • Request-path safety (httpguard.py, applied across proxy, aproxy, gateway): upstream-path validation (blocks @////.. host-injection SSRF), defensive Content-Length parsing, an 8 MiB body cap, and a bounded async connector.
  • Crash-resistance: compress_messages and ingest no longer raise on malformed-but-valid JSON (missing/non-string text, non-dict messages, bad JSONL lines) — they pass such input through untouched; the compress call in every proxy is additionally guarded so compression can never break a request.
  • Gateway: tenant labels are sanitized to a safe charset (no injection into accounting or the dashboard) and all HTML renderers (gateway, telemetry, ledger) escape interpolated values (stored-XSS fix).
  • Correctness: salience.protect() now falls back to the byte-exact original (never the stripped block) so a salient line is never silently dropped, and uses exact line membership; structured.fold leaves null-bearing records byte-exact (no null-vs-missing ambiguity); the Rust hot-path pins JSON key order to match the Python backend.

0.12.0

The Decision-Equivalence Risk Certificate (conformal risk control, distil conformal), salience protection (model-free frontier shifter), and the live head-to-head vs. the real LLMLingua-2 / Headroom packages. See BENCHMARKS.md.

0.9.0 – 0.11.0

Recoverable compression (distil_expand), the self-improving learning flywheel (distil learn), and the conformal certificate foundations.

0.2.0

Both sides of the bill, the proof pack, and the leapfrog tracks.

Added

  • Output compression — gated generation-side verbosity shaping + lossless output-on-re-entry digest + an A/B harness (answer-preservation gate); distil output-savings, distil proxy --shape-output.
  • Certified compression frontier — eval.py, distil eval: savings-vs- decision-equivalence curve where every point carries its certification verdict.
  • Self-distilling keep-model — online.py, distil online: learns from causal labels from your own traffic, retrains, promotes only if non-inferior.
  • Verifiable federated telemetry — telemetry.py, distil federated-leaderboard: HMAC-signed, content-free savings + verdict.
  • Async high-concurrency proxy — aproxy.py, distil proxy --async ([async]).
  • Rust hot-path core — rust/distil-core (PyO3), distil/native.py with a pure-Python parity fallback (transparent acceleration when built).
  • Managed gateway — gateway.py, distil gateway with a live per-tenant dashboard.
  • Real-trace ingestion — ingest.py, distil ingest (Anthropic + OpenAI shapes).
  • Performance benchmark — perf.py, distil perf (p50/p95).
  • Transformer keep-model — ONNX adapter + training pipeline (distil train-transformer); verified demo checkpoint on the release.
  • OpenAI role:"tool" messages now get the decision-aware reversible digest.

0.1.0

The first end-to-end cut: compression with a quality contract.

Added

  • Cache-aware cost engine (compress/cache_aware.py) — prices a multi-turn agent loop and proves naive recompression busts the prompt cache.
  • Risk-graded compression — Tier-0 provably-lossless transforms, Tier-1 reversible digest with retrieval handles, cache stabilization (schema canonicalization + volatile-field extraction), reject-if-bigger invariant.
  • Causal / counterfactual pruning (replay/ablation.py) — discovers context that never changes a decision.
  • Quality contract — TOST non-inferiority gate (certify/), decision-equivalence.
  • Multi-domain trajectory corpus (7 domains) + distil bench CI gate.
  • Auth-mode gating (policy.py) — lossless-only on subscription/OAuth.
  • Holdout A/B (certify/holdout.py) — savings with a bootstrap 95% CI.
  • Byte-fidelity gate (fidelity.py) — reversibility + append-only, distil verify.
  • Phase-7 building blocks — BM25 partial retrieval (retrieval.py), delta / append-only context (delta.py), keep-model codec (codec/), gist tool-schema caching (gist.py).
  • Runtime adapter (adapters/anthropic.py) — compress an Anthropic Messages request with no caller code change.
  • Billing-grade path — Anthropic count_tokens tokenizer and live AgentRunner (opt-in distil[live]).
  • Distributables — PyPI wheel/sdist, Docker image, single-file distil.pyz, CI + release workflows.