The scoreboard
Every context compressor claims a saving. This page scores them on what a user pays for: dollars per solved task, with every failed attempt's spend counted, and whether tasks still get solved. Each compressor is run by its own real, pinned package through the same harness and the same official grader. distil is one of the rows, not the referee's favourite: the same rules apply to it.
How it is measured, and why these estimators, is ADR 0024. Paired against plain on the same tasks; the interval on the difference is the harness's Wald interval with an exact McNemar test, at a pre-registered 5-point non-inferiority margin. A row that says pending run has no committed data, and nothing here estimates it. The runs that will fill it, with their costs, are in the run plan. To referee a compressor on your own traffic instead, see distil audit in the CLI reference.
SWE-bench Lite (official grader)
claude-sonnet-5-5 · effort medium · n = 298 tasks graded in every arm · run 2026-10-03. Source: benchmarks/results/swebench-outcome-300-medium.
| Compressor | $ per solved task | Solved (95% CI) | vs plain, pts [95% CI] | Total $ | Tokens in / out | Steps | Pinned version |
|---|---|---|---|---|---|---|---|
| plain (no compression) | $0.0387 | 218/298 · 73.2% (67.9%–77.9%) | baseline | $8.44 | 6,810,575 / 357,562 | 1360 | — |
| distil | $0.0393 | 213/298 · 71.5% (66.1%–76.3%) | -1.7 [-5.1, +1.7] · INCONCLUSIVE | $8.37 | 7,081,760 / 351,388 | 1427 | 1.56.3 |
| RTK | pending run | ||||||
| Headroom | not in this harness | ||||||
| Selective Context | pending run | ||||||
| Anthropic context editing | pending run | ||||||
SWE-bench Lite (official grader)
claude-sonnet-5-5 · effort low · n = 299 tasks graded in every arm · run 2026-10-02. Source: benchmarks/results/swebench-outcome-300-grepfix. Plain rows are reused from swebench-outcome-300 (same model, effort and tasks), as that run's report.md says.
| Compressor | $ per solved task | Solved (95% CI) | vs plain, pts [95% CI] | Total $ | Tokens in / out | Steps | Pinned version |
|---|---|---|---|---|---|---|---|
| plain (no compression) | $0.0341 | 210/299 · 70.2% (64.8%–75.1%) | baseline | $7.17 | 5,114,011 / 303,228 | 1251 | — |
| distil | $0.0351 | 204/299 · 68.2% (62.7%–73.2%) | -2.0 [-5.5, +1.4] · INCONCLUSIVE | $7.17 | 5,527,181 / 296,148 | 1290 | 1.56.3 |
| RTK | pending run | ||||||
| Headroom | not in this harness | ||||||
| Selective Context | pending run | ||||||
| Anthropic context editing | pending run | ||||||
Terminal-Bench 2 (cost_truth, neutral meter)
Pending. Not run yet: the pilot's canary preflights are committed, no analysed attempts..
| Compressor | $ per solved task | Solved (95% CI) | vs plain, pts [95% CI] | Total $ | Tokens in / out | Steps | Pinned version |
|---|---|---|---|---|---|---|---|
| control (no compression) | pending run | ||||||
| distil | pending run | ||||||
| RTK | pending run | ||||||
| Headroom | pending run | ||||||
| Selective Context | not in this harness | ||||||
| Anthropic context editing | not in this harness | ||||||
How to read it
- $ per solved task is the arm's total spend on the tasks graded in every arm, divided by the tasks it solved. A cheaper arm that solves fewer tasks can cost more per solved task.
- vs plain is the paired difference in tasks solved, in points, with its verdict at the 5-point margin. PILOT, INCONCLUSIVE and INFERIOR mean what they say; none of them is a pass.
- Short tasks. On these runs the median task took a handful of steps, so contexts stayed short. A compressor built for long sessions is not exercised much here; the Terminal-Bench table runs longer agent sessions.
- Every figure is regenerated by
scripts/build_scoreboard.pyfrom the run directories, and the machine-readable copy is scoreboard.json.