Metadata-Version: 2.4
Name: szl-retrieval-bench
Version: 0.3.0
Summary: Honest retrieval benchmark lane: BM25, TF-IDF dense, RRF hybrid, nDCG/Recall/P@k/R-precision/MRR/MAP, fairness gates, hash-chained receipts. Fail closed.
License-Expression: Apache-2.0
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# szl-retrieval-bench

Honest retrieval benchmark lane for the SZL estate. Sparse BM25 baseline,
classical TF-IDF dense lane, RRF hybrid fusion, and ranking metrics
(nDCG@k, Recall@k, P@k, R-precision, MRR, MAP) with fairness gates and
hash-chained receipts.

Doctrine: a benchmark that cannot run returns `BLOCKED` with a reason.
It never fabricates a metric. Run states: `MEASURED | BLOCKED | INVALID | FAILED`.

## Install and verify

```
pip install -e . pytest
python -m pytest tests/ -q                  # full offline suite
python -m szl_retrieval_bench.harness       # JSON demo; P@10 + R-precision in every lane/receipt
```

## Metric API

```python
from szl_retrieval_bench.metrics import precision_at_k, r_precision

ranked = ["d1", "d2", "d3", "d4", "d5"]
relevant = {"d1", "d3", "d9"}

assert precision_at_k(ranked, relevant, 2) == {
    "state": "MEASURED", "P@2": 0.5,
}
assert r_precision(ranked, relevant) == {
    "state": "MEASURED", "R_precision": 0.6667, "R": 3,
}
```

P@k always divides by `k`, including when fewer than `k` results were
returned. `k < 1` or a non-integer cutoff is `INVALID`; R-precision is
`INVALID` when there are no relevant documents. A valid ranking with no
overlap is `MEASURED` at `0.0` and remains present in per-query output,
aggregates, and receipts.

## Lanes

### Synthetic answer fixtures

`python -m szl_retrieval_bench.fixture` scores only caller-supplied synthetic
question/answer fixtures with `fixture_exact_match_v1`: NFC normalization,
casefolding, collapsed Unicode whitespace, and strict equality retaining
punctuation and signs. It does not run retrieval, generation, or a provider.
External performance remains `UNMEASURED`; this is not an official LongMemEval,
LoCoMo, or STATE-Bench evaluation. The existing ranking APIs are unchanged.

Use a JSON array of `{question_id, question, answer}` objects and JSONL
`{question_id, hypothesis}` predictions, with unique nonempty string IDs:

```sh
python -m szl_retrieval_bench.fixture --dataset fixture.questions.json \
  --dataset-variant fixture --predictions fixture.predictions.jsonl \
  --system synthetic --judge fixture_exact_match_v1 --out fixture-receipts
```

Missing predictions count as incorrect; empty datasets, empty prediction sets,
all-empty hypotheses, duplicate/unknown IDs, malformed JSON, and unsupported
judge/dataset modes raise errors. Inputs are capped at 1 MiB and 1,000 rows.
The legacy `exact` option aliases strict fixture equality, never substring matching.

Each unique output directory retains input/source snapshots, verdicts, and an
unsigned hash-bound receipt. `fixture.verify_receipt(path,
expected_evaluator_sha256=trusted_digest)` recomputes consistency and optionally
checks an independently obtained source digest. It never executes the snapshot.
Unsigned coordinated rewrites are not authenticated evidence. Normal imports
honestly label source identity as an unverified file snapshot. Output and input
parents must be private, caller-owned local directories: portable link checks
do not defeat concurrent hostile directory replacement. See
[integration provenance](docs/fixture-provenance.md) for coverage and limits.

### Retrieval ranking

- `bm25` - stdlib BM25 (k1=1.5, b=0.75). No downloads, no network.
- `tfidf-dense` — classical dense retrieval: L2-normalized TF-IDF vectors,
  cosine similarity. Real dense-vector math, deterministic, honestly labeled
  classical (not a neural embedding). A neural adapter plugs into the same
  `rank(query, doc_ids)` interface without touching the harness.
- `hybrid_rrf` — RRF(k=60) fusion of BM25 and a pluggable dense ranker.
  Calling it without a dense ranker returns `BLOCKED`, by design.
- `compare` — fairness gate: runs covering different query sets are `INVALID`.
- Receipts — every comparison can emit a SHA-256 hash-chained
  `UNSIGNED_HONEST` receipt. The receipt hashes the source harness declaration
  `szl-retrieval-bench` so downstream aggregators can reject relabelling.
  This is a hashed source declaration, not issuer authentication: the chain
  proves integrity + order of its contents, not who produced them.

Multi-vector / late-interaction (ColBERT-style) lives in a separate lane:
different memory profile, different fairness constraints. Not mixed here.

## Changelog highlights

- current: comparison receipts bind the `szl-retrieval-bench` harness name in
  the hashed payload so Wave 1 consolidation can verify source declaration
  without inferring identity from a filename or caller-supplied label.
- v0.3.0: P@k and R-precision added to the metric API, all measured lanes,
  demo output, aggregate leaderboards, and hash-chained comparison receipts;
  invalid cutoffs and undefined R-precision fail closed.
- v0.2.0: dense lane added; fixed a latent circular-reference bug in
  receipted comparisons (found by the new three-lane test before push —
  receipts now embed a snapshot, never a self-reference).
- v0.1.0: BM25, RRF, metrics, fairness gates, receipts, CI.

Apache-2.0 · Doctrine v11 · SZL Holdings
