Metadata-Version: 2.4
Name: pyterrier-tar
Version: 0.1.1
Summary: Replay-only technology-assisted-review stopping extensions for PyTerrier
Author: Aaron Fletcher
License-Expression: MPL-2.0
Project-URL: Upstream, https://github.com/terrier-org/pyterrier
Keywords: technology-assisted review,systematic review,stopping rules,high-recall retrieval,pyterrier
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Text Processing :: Indexing
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE.txt
License-File: THIRD_PARTY_NOTICES.md
Requires-Dist: ir-measures>=0.4.1
Requires-Dist: numpy
Requires-Dist: pandas
Requires-Dist: pyterrier<2,>=1.1.2
Requires-Dist: scipy
Provides-Extra: plot
Requires-Dist: matplotlib>=3.6; extra == "plot"
Provides-Extra: grl
Requires-Dist: gymnasium>=0.29; extra == "grl"
Requires-Dist: scikit-learn; extra == "grl"
Requires-Dist: stable-baselines3>=2.3; extra == "grl"
Provides-Extra: test
Requires-Dist: pytest>=7; extra == "test"
Requires-Dist: matplotlib; extra == "test"
Requires-Dist: scikit-learn; extra == "test"
Provides-Extra: tutorial
Requires-Dist: matplotlib; extra == "tutorial"
Requires-Dist: pyterrier[java]<2,>=1.1.2; extra == "tutorial"
Requires-Dist: scikit-learn; extra == "tutorial"
Dynamic: license-file

# pyterrier-tar

Replay-only technology-assisted-review (TAR) stopping extensions for [PyTerrier](https://github.com/terrier-org/pyterrier).

This is an independent package, not a PyTerrier fork.  It depends only on
documented public `pyterrier` APIs so that the dependency can be upgraded and
compatibility tested independently. It does not vendor PyTerrier code.

## Status

Alpha (0.1.1). Each stopping rule replays a recorded, labelled ranking and
reports where it would have stopped, why, and at what review cost. Rules whose
authors released code or results are checked against them, and the rest
against their published definitions; the Published Evidence section below
summarises every check and what it does not establish.

The package also registers the public CLEF eHealth TAR collections as
`tar:clef2017`, `tar:clef2018`, and `tar:clef2019`. These dataset adapters
hash-check the public archives and pinned released AutoTAR rankings before
using them for replay validation.

File names such as `BENCHMARKING.md` and paths under `examples/` refer to the
source distribution, which holds the full documents, the example scripts, and
the notebooks. The wheel that `pip install` uses holds only the package; the
Install section shows how to fetch the rest.

## Stopping rules

For resumable experiments, uncertainty and failure reports, ranking comparisons,
TARexp/CSV history import, and diagnostic plots, see the Benchmarking section
below.

Every rule is a PyTerrier transformer over a ranked, labelled trajectory
(`qid`, `rank`, `label`, in review order). `transform` returns the reviewed
prefix; `stop_report` gives the stop and its cost; `stop_trace` explains the
decision; `calculation_trace` shows every checkpoint the rule evaluated. A rule
sees labels only up to the checkpoint it is judging, which the test suite
enforces for all of them.

A rule that aims at a recall target takes `target_recall` as its first
argument, with no default, and a rule that needs the collection size takes
`collection_size` next: `tar.IPHyperbolic(.8, 3000)`, `tar.CMHHeuristic(.8, 3000)`.
Control-set rules take their screened frame first: `tar.QBCB(control, .8)`.

Rules and metrics that need the collection size either require
`collection_size` or accept it optionally. When it is optional (`SAFE`,
`BetaBinomial`, `stopping_metrics`) and omitted, the trajectory is taken to be
the whole collection, so pass it whenever the ranking is truncated.
`CLEFTARDataset.complete_ranking` appends a topic's unranked documents. A
`qid` is matched across frames as text, so `1` and `'1'` are the same query.

### Choosing one

```python
tar.catalogue()                        # every rule: family, promise, what it needs, source
tar.catalogue(promise='certificate')   # only the rules that carry a guarantee
```

Start from what you can supply. With **labels alone**, you can run any
trajectory rule or the checkpoint rules that also take a collection size. With
**calibrated probabilities** from your classifier, the estimation rules become
available. If you can **screen an independent random sample**, `QBCB` and
`TargetRecapture` are the only rules here that give a guarantee rather than an
estimate, and they charge that sample as review cost.

### What a rule promises

The four groups below say what a rule *reads*. What it *promises* is a
different split, and decides how it should be evaluated, because what a rule
promises decides what would falsify it. `tar.catalogue()` gives the promise of
each rule.

| Promise | Rules | What is promised | The check |
| --- | --- | --- | --- |
| Heuristic | `Kneedle`, `Budget`, `Rule2399`, `ReviewHalf`, `FixedRound`, `BatchPrecision`, `ConsecutiveIrrelevant`, `SAFE` | nothing about recall: they detect a flattening gain curve or exhaust a budget | the cost/recall trade-off, and excess cost over the oracle depth |
| Estimator | `CMHHeuristic`, `AnytimeCMH`, `PoissonPoint`, `IPHyperbolic`, `BaselineInclusionRate`, `Quant`, `QuantCI`, `AutoStop`, `SCAL`, `Chao`, `BetaBinomial`, the `EVPI` family, `GRLStop` | recall at or above the target **if its model holds** | `reliability_summary`: the share of topics reaching the target, with a Wilson interval, against the nominal level |
| Certificate | `QBCB`, `TargetRecapture` | recall at or above the target with probability `1 - alpha` **over the draw of the control sample** | `coverage_summary`: the share of repeated draws that certified the target, for one review |

The unit of repetition is the difference that matters. An estimator's promise
is about topics, so it is measured across them. A certificate's promise is
about the sample it drew, so it is measured by drawing again: coverage across
replications of one review. Measuring a certificate across topics answers a
question nobody asked.

Costs are comparable only when the control screening is charged. `stop_report`
and `stopping_metrics` already count it once; a certificate that costs more
than a heuristic is buying a guarantee the heuristic never offers.
`EVALUATION.md` is the same guidance as a single page.

### Trajectory rules — labels only

| Rule | Stops when | Source |
| --- | --- | --- |
| `Kneedle` | At BMI checkpoints it locates the knee of the gain curve, then compares the pre-knee to post-knee slope ratio against `156 - min(relevant at knee, 150)`. `knee_distance='absolute'` (default) matches the released implementations; `'signed'` is Satopää et al.'s Kneedle. | [Cormack and Grossman (2016a)](https://doi.org/10.1145/2911451.2911510) |
| `Budget` | Three quarters of the collection is reviewed, or `n >= 10N/r` with a knee slope ratio of at least 6. | [Cormack and Grossman (2016a)](https://doi.org/10.1145/2911451.2911510) |
| `Rule2399` | Reviewed documents reach `2399 + 1.2 x relevant found`. | [Cormack and Grossman (2016b)](https://doi.org/10.1145/2983323.2983776) |
| `ReviewHalf` | Half the known collection is reviewed. | [Yang and Lewis (2022), TARexp](https://doi.org/10.1145/3477495.3531663) |
| `FixedRound` | A set number of review batches is done. | [Yang and Lewis (2022), TARexp](https://doi.org/10.1145/3477495.3531663) |
| `BatchPrecision` | `patience` consecutive batches have precision at or below the cutoff. | [Yang, Lewis, and Frieder (2021)](https://doi.org/10.1145/3469096.3469873) |
| `ConsecutiveIrrelevant` | A run of `count` non-relevant documents completes. | common heuristic |
| `SAFE` | Every supplied key paper is found, at least twice the seed positives and `min_fraction` of the collection are reviewed, and the last `consecutive` documents are non-relevant. | [Boetje and van de Schoot (2024)](https://doi.org/10.1186/s13643-024-02502-7) |
| `Oracle` | The prefix first reaches the target recall, using complete labels. An offline lower bound, never deployable. | — |

### Checkpoint rules — a statistical test at fixed intervals

| Rule | Stops when | Source |
| --- | --- | --- |
| `CMHHeuristic` | The biased-urn (BUSCAR) test rejects "recall is still below target" at `alpha`. | [Callaghan and Müller-Hansen (2020)](https://doi.org/10.1186/s13643-020-01521-4) |
| `AnytimeCMH` | The same test with an `alpha / (k(k+1))` spending schedule across checkpoints, so the levels sum to `alpha`. Conditional on the urn p-values being calibrated for the recorded design. | original to this package (A. Fletcher) |
| `PoissonPoint` (IP-P) | A power-law rate fitted to window relevance yields a Poisson upper bound on the documents still unfound, and `found >= target x (found + bound)`; also when the later windows contain nothing relevant. | [Stevenson and Bin-Hezam (2023)](https://doi.org/10.1145/3631990) |
| `IPHyperbolic` (IP-H) | As IP-P with a hyperbolic rate. `tail='released'` (default) reproduces the released code's expected-remaining formula, which understates it by `(1-b)^2`; `tail='model'` integrates the fitted rate. | [Stevenson and Bin-Hezam (2023)](https://doi.org/10.1145/3631990) |

### Control-set rules — an independently screened sample

| Rule | Stops when | Source |
| --- | --- | --- |
| `QBCB` | The `J`th positive control appears in the ranking, where `J` is the smallest rank whose binomial bound certifies the target. | [Lewis, Yang, and Frieder (2021)](https://doi.org/10.1145/3459637.3482415) |
| `TargetRecapture` | All `k = ceil(-ln(1-confidence)/(1-target))` sampled targets have been reviewed. | [Cormack and Grossman (2016a)](https://doi.org/10.1145/2911451.2911510) |
| `BaselineInclusionRate` | Relevant documents found reach `target x` the prevalence estimated from a random pilot. It never fires when the pilot finds nothing. | pilot-prevalence heuristic |
| `sample_control` | Helper: draws an unlabelled worklist to screen independently. Control screening is charged in `stop_report` and `stopping_metrics`. | — |

### Estimation rules — recorded probabilities or sampling metadata

| Rule | Stops when | Source |
| --- | --- | --- |
| `Quant`, `QuantCI` | Estimated recall from calibrated probabilities reaches the target; `QuantCI` first subtracts `nstd` standard deviations. Pass `score_snapshots` to replay per-round model scores instead of one fixed column. | [Yang, Lewis, and Frieder (2021)](https://doi.org/10.1145/3469096.3469873) |
| `AutoStop` | A Horvitz-Thompson total over a fixed with-replacement sample certifies the target, under `loose`, `strict_v1`, or `strict_v2`; `strict_v2` needs `collection_size`. Not the adaptive procedure. | [Li and Kanoulas (2020)](https://doi.org/10.1145/3411755) |
| `SCAL` | The prefix's Horvitz-Thompson estimate reaches the target share of the estimate over the whole recorded sample. The sample past the cutoff is charged as review cost; pass `control_documents()` to `stopping_metrics` so it counts towards recall too. | [Cormack and Grossman (2016b)](https://doi.org/10.1145/2983323.2983776) |
| `Chao` | The Chao1 estimate of unseen relevant documents satisfies `found >= target x (found + unseen)`. Simpler than the paper's multi-model estimators. | [Bron et al. (2025)](https://doi.org/10.1145/3724116) |
| `BetaBinomial` | The beta-binomial posterior that few enough relevant documents remain exceeds `confidence`. Assumes the unreviewed tail is exchangeable with the reviewed prefix, so it errs late on a priority ranking. | original to this package (A. Fletcher) |
| `EVPI`, `EVPIGreedy`, `EVPISmooth`, `EVPIBatch` | The expected value of perfect information about the next document (or window) falls to the screening cost. The document that triggers the stop is included in the reviewed prefix. | original to this package (A. Fletcher) |

`AnytimeCMH`, `BetaBinomial`, and the `EVPI` family are original to this
package rather than replays of published methods, so they carry no parity
claim: they are covered by formula tests and the shared leakage and trace
checks only.

### Learned rule

`GRLStop` trains a PPO policy over 100 review windows, where the state is the
relevance rate in each reviewed window, a logistic-regression prediction for
each unreviewed one, the current window, and the target recall. It needs the
`grl` extra. Source: [Bin-Hezam and Stevenson (2025)](https://doi.org/10.1145/3726302.3729879).

### Datasets and helpers

`CLEFTARDataset` backs `pt.get_dataset('tar:clef2017')` and its 2018/2019
siblings: hash-pinned public archives, one candidate pool per topic,
`complete_ranking()` to append unretrieved documents, and `label_ranking()` to
attach qrels only after a ranking exists. Data is downloaded into PyTerrier's
cache, never this package.

```python
import pyterrier as pt

dataset = pt.get_dataset('tar:clef2017')

# Replay the pinned, released AutoTAR ranking of one topic.
ranked = dataset.get_results('autotar').query("qid == 'CD008081'")
trajectory = dataset.label_ranking(dataset.complete_ranking(ranked, 'CD008081'))
reviewed = pt.tar.Kneedle().transform(trajectory)
```

To rank a topic yourself, build an index for it, because each
systematic-review topic has its own candidate pool (this needs the `tutorial`
extra):

```python
from tempfile import TemporaryDirectory

import pandas as pd

topic = dataset.get_topics().query("qid == 'CD008081'")
docs = pd.DataFrame(dataset.get_corpus_iter('CD008081'))
with TemporaryDirectory() as index_dir:
    indexref = pt.index.IterDictIndexer(index_dir, meta={'docno': 64}).index(docs.to_dict('records'))
    bm25 = pt.terrier.Retriever(indexref, wmodel='BM25', num_results=len(docs))
    # The topic queries are Boolean PubMed searches; tokenise() reduces them to plain terms.
    ranked = (pt.rewrite.tokenise() >> bm25).transform(topic)
trajectory = dataset.label_ranking(dataset.complete_ranking(ranked, 'CD008081'))
```

Labels are joined only after a ranking has been created; they must never be
made available to a ranker or a live stopping decision.

### Metrics

`stopping_metrics` reports recall, cost, review fraction, reliability, oracle
depth, and the CLEF `loss_e`/`loss_r`/`loss_er` components, charging any
separately screened controls once. `stopping_frontier` summarises several
targets, and `reliability` and `review_fraction` are ir-measures metrics for
`pt.Experiment`.

`reliability_summary` gives the share of topics reaching the target with a
Wilson interval, which is the check for an estimator. `coverage_summary` gives
the share of repeated control draws that certified it, which is the check for a
certificate.

### Adding a rule

One file per approach, inside the folder for what the rule reads:
`rules/trajectory/` for labels only, `rules/checkpoint/` for a test at fixed
checkpoints, `rules/control/` for an independently screened sample, and
`rules/estimation/` for recorded probabilities or sampling metadata.

```python
# src/pyterrier_tar/rules/trajectory/my_rule.py
from .._base import _TrajectoryStoppingRule, _batch_positions


class MyRule(_TrajectoryStoppingRule):
    """Stops once a batch holds no relevant documents."""

    _equation = 'fire if the batch ending at n has no relevant document'

    def __init__(self, batch_size: int = 200, initial_documents: int = 1):
        if batch_size < 1 or initial_documents < 1:
            raise ValueError('batch sizes must be positive')
        self.batch_size = batch_size
        self.initial_documents = initial_documents

    def _checkpoint_rows(self, ranked, labels):
        for stop in _batch_positions(len(labels), self.batch_size, self.initial_documents):
            batch = labels[max(0, stop - self.batch_size):stop]
            yield {'index': stop, 'relevant_in_batch': int((batch > 0).sum()),
                   'fired': not (batch > 0).any()}
```

`_checkpoint_rows` yields one dict per checkpoint, with at least `index` and
`fired`; the base class turns it into the stop, `stop_report`, `stop_trace`, and
`calculation_trace`, so those cannot disagree. A rule that is not
checkpoint-shaped implements `_stop_position(labels)` instead, and one that
needs other columns implements `_stop_position_for_results(ranked, labels)`.
Add `_trace_details` for the columns `stop_trace` should carry.

The row for checkpoint `k` must use only `labels[:k]`. Export the class from
the group's `__init__.py`, from `rules/__init__.py`, and from
`pyterrier_tar/__init__.py`; the shared tests then pick it up, including the
leakage check and the report/trace consistency check. New rules are expected to
come with a test that re-derives the stop from the rule's own definition, as
`tests/test_rule_validity.py` does for the others.

## Scope and safety boundary

`pyterrier-tar` replays recorded, labelled trajectories.  It is not a live active-learning or screening controller, and it does not certify target recall in deployment.  Labels beyond the reviewed prefix must not be exposed to a stopping rule.  Separately sampled controls must be independently screened and accounted for in review cost.

## Install

```bash
python -m pip install pyterrier-tar
python -m pip install 'pyterrier-tar[grl]'
```

The `grl` extra installs `gymnasium`, `scikit-learn`, and `stable-baselines3`
for `GRLStop`; `plot` adds matplotlib for the figures; `tutorial` adds
PyTerrier's Java support and matplotlib for the notebooks that build Terrier
indexes.

The documents this README names, the example scripts, the notebooks, and the
tests are in the source distribution, not the wheel:

```bash
python -m pip download pyterrier-tar --no-deps --no-binary :all:
tar -xzf pyterrier_tar-*.tar.gz
```

The unpacked directory holds `BENCHMARKING.md`, `EVALUATION.md`,
`PUBLISHED_EVIDENCE.md`, `GAPS.md`, `RESULTS.md`, `RESULTS_UNCERTAINTY.md`,
`CHANGELOG.md`, `examples/`, and `tests/`. The same archive is on the
[PyPI files page](https://pypi.org/project/pyterrier-tar/#files).

## Quick start

```python
import numpy as np
import pandas as pd
from pyterrier import tar

# A recorded review: ranked documents, with the label revealed as each is read.
# Relevance thins out down the ranking, as it does in a real screening run.
rng = np.random.default_rng(0)
size = 6000
prevalence = .7 * np.exp(-np.arange(size) / 300) + .002
trajectory = pd.DataFrame({
    'qid': 'CD008081',
    'docno': [f'd{position}' for position in range(size)],
    'rank': range(size),
    'label': (rng.random(size) < prevalence).astype(int),   # 1 for relevant
})

rule = tar.Kneedle()
reviewed = rule.transform(trajectory)          # the prefix the rule would have read
report = rule.stop_report(trajectory)          # stop, fired, review cost
print(rule.stop_trace(trajectory).iloc[0].reason)
print(rule.calculation_trace(trajectory, index=int(report.stop.iloc[0])).tail())

qrels = trajectory.loc[trajectory.label > 0, ['qid', 'docno', 'label']]
print(tar.stopping_metrics(trajectory, reviewed, qrels, target_recall=.95, report=report))
```

On this trajectory Kneedle stops after 2,001 of 6,000 documents, having found
94.7% of the relevant ones; `stop_trace` says the knee slope ratio reached its
dynamic threshold, and `calculation_trace` shows the ratio at every checkpoint
it tested.

## Benchmarking

The experiment API runs recorded histories, saves per-topic results, and
reports uncertainty and failures. The input ranking is fixed before labels are
attached; this workflow does not train a ranker or select a stopping rule.

```python
from pathlib import Path
import pyterrier_tar as tar

# trajectory has qid, docno, rank, label and covers the full collection.
data = tar.ReplayDataset("my-collection", "recorded-ranker", trajectory)
rules = [
    tar.RuleSpec("IP-P", lambda c: tar.PoissonPoint(c.target, c.collection_size)),
    tar.RuleSpec("50 irrelevant", lambda c: tar.ConsecutiveIrrelevant(50),
                 config={"count": 50}),
]
output = Path("artifacts/my-experiment")
results = tar.run_experiment([data], rules, targets=[.8, .9, .95],
                             seeds=[0], output=output)
summary = tar.experiment_summary(results)
summary.to_csv(output / "summary.csv", index=False)
(output / "REPORT.md").write_text(tar.experiment_report(results))
```

* **Rules.** A factory receives `RuleContext(trajectory, target, seed,
  collection_size)` and returns a fresh rule. Built-in rules are classified
  from the catalogue; a custom class needs `promise=` (`heuristic`,
  `estimator`, `certificate`, or `offline bound`), which is metadata, not a
  verified guarantee. A factory that samples must draw from `context.rng()`,
  which is keyed by topic, target and seed; `default_rng(context.seed)` alone
  repeats one stream for every topic and target.
* **Repetition.** `repetitions="once"` uses the first seed and `"seeds"` uses
  every seed. Certificates default to `"seeds"` and other rules to `"once"`.
* **Resume and provenance.** Completed tasks are written atomically under
  `tasks/`, and repeating the call resumes them. `manifest.json` binds the
  parameters, input-frame hashes, package source hash and dependency versions;
  changing the manifest requires a new directory. Only one writer may use a
  directory.
* **Partial histories.** Omitting `qrels` declares the trajectory to be the
  entire collection, and the runner cannot infer unseen relevant documents.
  For a partial history, pass full-pool qrels including negatives.
* **Uncertainty and failures.** `experiment_summary` gives target attainment
  with Wilson 95% intervals, mean review fraction with bootstrap intervals,
  minimum recall, recall shortfalls, and counts of firings, full-review
  fallbacks, exhausted partial histories and zero-relevant topics.
  `worst_topics` keeps the topic, target and seed of individual failures. For a
  certificate, each row is one topic across independent control draws, and
  certification frequency is separate from reaching the target: reviewing
  everything can reach the target without producing a certificate. The
  intervals are descriptive and do not adjust for selecting among methods or
  targets.
* **Baseline rankings.** `bm25_ranking` and `random_ranking` are generated
  before qrels are joined, and `complete_rankings(pool, ranking)` appends
  unranked documents, so rankers are compared on the same pool. `bm25_ranking`
  is a specified in-memory baseline and claims no Terrier scoring parity.
* **Review histories.** `read_review_history("events.csv", scores="scores.csv")`
  reads portable CSV events (`qid,docno,rank,label,round`); `import_tarexp`
  and `read_tarexp` read TARexp ledgers. Native TARexp checkpoints are gzipped
  Python pickles and may execute code, so `read_tarexp` requires
  `trusted=True`: only load checkpoints you trust. TARexp records batch
  membership, not within-batch review order, so a document-level stop inside
  an imported batch is not a reconstructed historical decision.
* **Figures.** With the `plot` extra, `figure_recall_effort`,
  `figure_target_misses`, `figure_control_cost`, `figure_stopping_outcomes`
  and `figure_stopping_diagnostic` return figures 180 mm wide, and
  `save_figure(fig, "figures/name")` writes PDF, SVG and 600-dpi PNG.

`BENCHMARKING.md` has the rest: what the manifest records, ranking comparisons
on CLEF with `examples/benchmark_rankings.py`, the history formats, and what
each figure shows.

## Development

From the unpacked source distribution:

```bash
python -m pip install -e '.[test,grl,tutorial]'
pytest
```

`test` runs the suite, `grl` adds the GRLStop tests, and `tutorial` the
notebooks. On a CUDA-free machine, `UV_TORCH_BACKEND=cpu` keeps `grl` from
pulling the GPU wheels.

## Results on CLEF

Every rule replays the released AutoTAR ranking of each topic in CLEF 2017,
2018, and 2019 (91 topics with a relevant document: 30 from 2017, 30 from 2018,
31 from 2019), completed with the pool's unranked documents. Review is the
share of each topic's pool screened, averaged over topics; for certificates it
includes the control sample.

**These numbers describe one ranker.** A rule that reads the ranking, which is
every rule here except the certificates, can do much worse when the ranking is
weaker. Rules are grouped by what they promise, so compare within a group, not
across groups.

An estimator promises the target if its model holds. Reliability is the share
of topics that reached the target; the rule was asked for 95% confidence, so
95% reliability is what it promises. The Oracle reads every label, so no
deployable rule can match it; it marks the least review each target could have
needed.

| Rule | Reliability at 0.8 / 0.9 / 0.95 | Review at 0.8 / 0.9 / 0.95 |
| --- | --- | --- |
| CMHHeuristic | 100.0% / 100.0% / 100.0% | 47.6% / 60.8% / 72.1% |
| AnytimeCMH | 100.0% / 100.0% / 100.0% | 63.8% / 78.1% / 90.1% |
| PoissonPoint (IP-P) | 100.0% / 100.0% / 100.0% | 27.7% / 28.6% / 28.9% |
| IPHyperbolic (IP-H) | 93.4% / 90.1% / 89.0% | 16.1% / 16.8% / 18.1% |
| IPHyperbolic (IP-H), tail='model' | 100.0% / 97.8% / 97.8% | 17.6% / 18.5% / 20.0% |
| BetaBinomial | 100.0% / 100.0% / 100.0% | 90.3% / 96.5% / 99.0% |
| Oracle (offline bound) | 100.0% / 100.0% / 100.0% | 5.1% / 6.5% / 7.8% |

A heuristic takes no target and promises nothing about recall. It stops in the
same place whatever the target; the columns show how often that happened to
reach each one.

| Rule | Review | Reached 0.8 / 0.9 / 0.95 |
| --- | --- | --- |
| Kneedle | 54.7% | 100.0% / 100.0% / 100.0% |
| Budget | 43.8% | 100.0% / 100.0% / 100.0% |
| Rule2399 | 72.3% | 100.0% / 100.0% / 100.0% |
| ReviewHalf | 58.8% | 100.0% / 100.0% / 100.0% |
| BatchPrecision | 31.5% | 100.0% / 98.9% / 96.7% |
| ConsecutiveIrrelevant | 20.8% | 100.0% / 97.8% / 95.6% |
| SAFE (no seed positives) | 18.0% | 100.0% / 100.0% / 94.5% |

A certificate promises the target with 95% probability over the draw of its
control sample, so it is measured by coverage: the share of 20 simulated
samples per topic that reached the target. A topic with too few relevant
documents cannot be certified, and the rule then reviews everything.

| Rule | Target | Coverage | Topics it could certify | Review (of which control) |
| --- | --- | --- | --- | --- |
| QBCB | 0.8 | 98.0% | 71 of 91 | 48.5% (42.3%) |
| QBCB | 0.9 | 98.2% | 60 of 91 | 65.6% (60.0%) |
| QBCB | 0.95 | 99.1% | 39 of 91 | 83.5% (79.0%) |
| TargetRecapture | 0.8 | 98.5% | 71 of 91 | 50.0% (43.8%) |
| TargetRecapture | 0.9 | 98.6% | 59 of 91 | 66.5% (60.9%) |
| TargetRecapture | 0.95 | 99.2% | 38 of 91 | 83.8% (79.4%) |

`Quant`, `QuantCI`, `AutoStop`, `SCAL`, `Chao`, the `EVPI` family, `GRLStop`,
`FixedRound`, and `BaselineInclusionRate` are not in these tables: each needs
an input the CLEF release does not contain, a trained policy, or a parameter
whose value would be an arbitrary choice.

These tables are copied from `RESULTS.md`, which adds the breakdown by CLEF
year. `RESULTS_UNCERTAINTY.md` gives the intervals behind each row and the
largest individual shortfalls. `examples/results_table.py` regenerates both.

## Published Evidence

A passing check does not turn a heuristic rule into a deployment recall
certificate. Each row says what was compared and how far the claim goes.

| Rule or method | Checked against | What the check establishes |
| --- | --- | --- |
| `stopping_metrics` | the official CLEF `tar_eval.py` on its two sample runs, all 20 topics | recall, total cost, and the three loss fields match at 3 decimals; no claim for the CLEF metrics it does not expose, such as NCG, AP, or WSS |
| `IPHyperbolic` (IP-H) | the authors' released rankings and results for CLEF 2017--19 | strict parity: recall, cost, reliability, `loss_er`, and relative error in all nine settings at 3 decimals |
| `PoissonPoint` (IP-P) | the same release, CLEF 2017 | strict parity: every published IP-P and IP-H aggregate at 3 decimals |
| `Oracle` | the released point-process baseline table | all nine CLEF 2017--19 oracle rows match |
| `Kneedle` | Sneyd and Stevenson's source on the CLEF 2017 Waterloo run | the source's total effort (86,243) and 0.998333 mean recall, with `min_documents=1, batch_size=1`; the default fixed-200 replay is a compatibility mapping |
| `CMHHeuristic` | `buscarpy`, pinned | 216 p-values and five retrospective stop positions match; source-definition parity, not a paper reproduction |
| `QBCB` | Lewis, Yang, and Frieder's Table 1 | every published `r -> J` control rank; no numerical paper parity, because the paper's inputs were not released |
| `AutoStop` | the released AutoStop source, CLEF topic CD008081 | the `loose`, `strict_v1`, and `strict_v2` estimators on a fixed sampling distribution; not the adaptive procedure |
| `Quant`, `QuantCI`, `BatchPrecision`, `Rule2399`, `ReviewHalf`, `FixedRound` | TARexp's definitions | definitions follow TARexp and Quant's variance mirrors it; no paper result is asserted |
| `Budget`, `SCAL`, `Chao`, `TargetRecapture`, `BaselineInclusionRate` | their published definitions | definition checks only; `Chao` is the classic Chao1, simpler than the paper's estimators |
| `GRLStop` | the released GRLStop source | compared line by line; no trained policies or result tables were released, so there is nothing numerical to reproduce |
| `AnytimeCMH`, `BetaBinomial`, the `EVPI` family | nothing external | original to this package: formula tests and the shared leakage and trace checks only |

Every rule that decides from a reviewed prefix is also checked for leakage:
changing the labels after its stop, or after a trace index, never changes it.

`PUBLISHED_EVIDENCE.md` is the full ledger, with the pinned revisions and
hashes behind each row. The examples under `examples/` are intentionally
separate from CI when they require network access, large public archives, or
locally held trajectories.

For real-data or external-source smoke checks:

```bash
python examples/published_cmh_buscar.py
python examples/published_clef_tar_eval_metrics.py
python examples/published_ip_h_clef.py
python examples/published_kneedle_clef2017.py
python examples/published_point_process_clef2017.py
python examples/published_autostop_clef2017.py
PYTHONPATH=examples python examples/published_baselines_clef.py
```

Run them from a scratch directory: each caches its inputs in `data/` relative
to the working directory, about 1 GB in total.

`examples/notebooks/parity_published.ipynb` runs all of them and shows one
parity table. Every notebook is either a `demo_` (how the package behaves) or a
`parity_` (a published result reproduced):

| Notebook | Purpose |
| --- | --- |
| `demo_stopping_rules` | Every public stopping rule on a small synthetic trajectory: its decision, trace, and replay-only boundary. |
| `demo_trajectory_rules` | Why each trajectory-only rule fires, drawn: revealed labels, checkpoints, the trigger quantity, and the stop. |
| `demo_workflows` | Two worked workflows on real collections: a CLEF topic from index to stop, and control-set stopping on Vaswani with its screening cost. |
| `demo_fresh_rankers_stopping` | Build fresh BM25/TF-IDF rankings, run editable stopping methods in memory, and display publication figures. |
| `demo_multitopic_rankers_stopping` | The same across ten topics by default or a whole CLEF year, with all-topic outcomes and a paired diagnostic. |
| `demo_publication_figures` | Load and verify saved benchmark results, then display and export the publication figures. |
| `parity_published` | One table for every published claim: runs each networked parity script and asserts its result. |
| `parity_archives` | Four audits of other groups' released results: Repke et al. (2026), König et al. (2024), Chao et al. (2024), and the SYNERGY v2 Kneedle screen. Each needs archives you download, and skips when they are absent. |

`published_kneedle_clef2017.py` checks Kneedle's released 86,243 effort. The
same seven scripts run weekly in CI's `published-parity` job. The notebooks that
build Terrier indexes need the `tutorial` extra, which installs PyTerrier's
Java support.

Only the metrics asserted by each evidence check should be treated as validated.
For example, the CLEF metric checker asserts only the official recall, total
cost, and loss fields exposed by `stopping_metrics()`, while the IP-H CLEF
checker asserts recall, cost, reliability, loss, and relative-error parity
across all nine CLEF 2017--19 settings.

## Known gaps

What this package does not implement or cannot reproduce, and why. It replays
recorded reviews, so nothing here establishes how a rule behaves when it
steers a live one.

These published results need data the authors did not release. Nothing in the
package should claim them.

| Result | Missing input |
| --- | --- |
| QBCB's RCV1 figures | the 20% subset, seeds, active-learning trajectories, and control-sample identities |
| Yang et al. (2021) heuristics on RCV1 | the saved review histories |
| Adaptive AutoStop | the per-draw sampling distributions, which a trajectory cannot reconstruct |
| Chao et al. Table 7 integer displays | the analysis and formatting code; raw aggregates do not round to the printed integers (57.566 prints as 57) |
| The `Knee` rows of the point-process baseline table | the 2018 qrels are rebuilt in that repository from a PID list it does not ship; `published_baselines_clef.py` reports the difference instead of asserting it |
| `TM` and `TM-adapted` rows of the same table | the released target method draws its target set with an unseeded `random.choice` |
| RLStop and GRLStop paper results | no trained policies or result tables were released |

These methods are named but not implemented.

| Name | State |
| --- | --- |
| Score-distribution stopping (Hollmann and Eickhoff) | not implemented; `QuantCI` over calibrated probabilities is the nearest rule |
| The adapted target method of the point-process papers | not implemented; `TargetRecapture` is Cormack and Grossman's original target method |
| TARexp-style Knee and Budget | TARexp takes the maximum slope ratio over every round split, measured in rounds. This package follows Cormack and Grossman's knee. An opt-in mode would be needed for TARexp-equivalent numbers. |
| Adaptive AutoStop, a live S-CAL controller | out of scope: this package replays recorded trajectories and does not select documents |

And these are limits of the engineering.

| Gap | Notes |
| --- | --- |
| CLEF archives come from one mirror | `npai.science.uu.nl`, hash-pinned. The hashes protect the content, not its availability |
| 9 of 119 SYNERGY datasets will not compose | missing source files or server errors at the pinned commit |
| Parity runs weekly, archive audits by hand | the `published-parity` job downloads about 1 GB; the archive notebooks need artifacts you supply |

These tables are copied from `GAPS.md`. New findings about limits belong there
rather than in a commit message.

## Compatibility policy

The supported PyTerrier range is declared in `pyproject.toml`.  New code must import only public `pyterrier` names; imports from private modules such as `pyterrier._ops` are prohibited.  Each supported PyTerrier release will receive a focused compatibility test before its range is widened.

## License and notices

This source is licensed under the
[Mozilla Public License 2.0](https://www.mozilla.org/MPL/2.0/). `LICENSE.txt`
and `THIRD_PARTY_NOTICES.md` are installed with the package, in the
`licenses` folder of its `dist-info` directory.

The package contains no third-party source code. It was first written in a
fork of PyTerrier and extracted into this package; it uses PyTerrier only
through PyTerrier's public API. The stopping rules implement published methods
from their papers. Where a rule reproduces a released configuration, such as
`IPHyperbolic`'s released tail formula, it is an implementation of the
published method, not a copy of the authors' source code.
