compression with a quality contract

The scoreboard

Every context compressor claims a saving. This page scores them on what a user pays for: dollars per solved task, with every failed attempt's spend counted, and whether tasks still get solved. Each compressor is run by its own real, pinned package through the same harness and the same official grader. distil is one of the rows, not the referee's favourite: the same rules apply to it.

How it is measured, and why these estimators, is ADR 0024. Paired against plain on the same tasks; the interval on the difference is the harness's Wald interval with an exact McNemar test, at a pre-registered 5-point non-inferiority margin. A row that says pending run has no committed data, and nothing here estimates it. The runs that will fill it, with their costs, are in the run plan. To referee a compressor on your own traffic instead, see distil audit in the CLI reference.

SWE-bench Lite (official grader)

claude-sonnet-5-5 · effort medium · n = 298 tasks graded in every arm · run 2026-10-03. Source: benchmarks/results/swebench-outcome-300-medium.

Compressor$ per solved taskSolved (95% CI)vs plain, pts [95% CI]Total $Tokens in / outStepsPinned version
plain (no compression)$0.0387218/298 · 73.2% (67.9%–77.9%)baseline$8.446,810,575 / 357,5621360—
distil$0.0393213/298 · 71.5% (66.1%–76.3%)-1.7 [-5.1, +1.7] · INCONCLUSIVE$8.377,081,760 / 351,38814271.56.3
RTKpending run
Headroomnot in this harness
Selective Contextpending run
Anthropic context editingpending run

SWE-bench Lite (official grader)

claude-sonnet-5-5 · effort low · n = 299 tasks graded in every arm · run 2026-10-02. Source: benchmarks/results/swebench-outcome-300-grepfix. Plain rows are reused from swebench-outcome-300 (same model, effort and tasks), as that run's report.md says.

Compressor$ per solved taskSolved (95% CI)vs plain, pts [95% CI]Total $Tokens in / outStepsPinned version
plain (no compression)$0.0341210/299 · 70.2% (64.8%–75.1%)baseline$7.175,114,011 / 303,2281251—
distil$0.0351204/299 · 68.2% (62.7%–73.2%)-2.0 [-5.5, +1.4] · INCONCLUSIVE$7.175,527,181 / 296,14812901.56.3
RTKpending run
Headroomnot in this harness
Selective Contextpending run
Anthropic context editingpending run

Terminal-Bench 2 (cost_truth, neutral meter)

Pending. Not run yet: the pilot's canary preflights are committed, no analysed attempts..

Compressor$ per solved taskSolved (95% CI)vs plain, pts [95% CI]Total $Tokens in / outStepsPinned version
control (no compression)pending run
distilpending run
RTKpending run
Headroompending run
Selective Contextnot in this harness
Anthropic context editingnot in this harness

How to read it