Metadata-Version: 2.4
Name: judgemetry
Version: 0.1.0
Summary: A reliability lab for LLM judges: measure position bias, verbosity bias, and calibration against human labels — then recalibrate. Pure stdlib; no API keys needed for the demo.
Project-URL: Homepage, https://github.com/sedai77/judgemetry-llm-judge-reliability
Project-URL: Documentation, https://github.com/sedai77/judgemetry-llm-judge-reliability#readme
Project-URL: Changelog, https://github.com/sedai77/judgemetry-llm-judge-reliability/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/sedai77/judgemetry-llm-judge-reliability/issues
Author: Judgemetry contributors
License-Expression: MIT
License-File: LICENSE
Keywords: benchmark,calibration,evals,evaluation,isotonic-regression,llm,llm-as-judge,mt-bench,position-bias,reliability
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.10
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.77; extra == 'anthropic'
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Description-Content-Type: text/markdown

# Judgemetry

> A reliability lab for LLM judges. Measure position bias, verbosity bias, and
> calibration against human labels — then fix the calibration.

[![CI](https://github.com/sedai77/judgemetry-llm-judge-reliability/actions/workflows/ci.yml/badge.svg)](https://github.com/sedai77/judgemetry-llm-judge-reliability/actions/workflows/ci.yml)
[![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue.svg)](pyproject.toml)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)

Model-graded evals are everywhere: an LLM-as-judge picks the better of two
answers, and that pairwise preference drives leaderboards, regression gates,
and RLHF-adjacent pipelines. The judge itself is almost never validated.
Judges systematically prefer whichever answer they read first (**position
bias**), reward longer answers beyond what humans do (**verbosity bias**),
favor their own model family, and report confidences that don't match their
accuracy (**eval calibration**). Judgemetry treats the judge like a lab
instrument: it measures each failure mode against human preference labels,
puts a bootstrap confidence interval on every number, recalibrates the
confidence channel, and hands you a one-file HTML report card with letter
grades — so "we use an LLM judge" can come with an error bar.

**Zero runtime dependencies.** Pure Python standard library — the bootstrap,
Cohen's kappa, ECE, isotonic regression, and even the SVG charts are
hand-rolled. The one optional extra (`anthropic`) is imported lazily and only
needed for live runs. Python 3.10+.

## 60-second start (no keys, no network)

```bash
pip install git+https://github.com/sedai77/judgemetry-llm-judge-reliability
judgemetry demo
```

(Not on PyPI yet — install from git until the first release is published;
`pip install judgemetry` will work after that.)

That runs the **entire** pipeline — dataset, double-pass judging with swapped
answer positions, metrics with bootstrap CIs, isotonic recalibration, HTML
report — on a bundled synthetic fixture with *planted* biases, judged by a
deterministic simulated judge. No API key, no network call, no other package.
It prints the report card and writes `report.html`. Actual output (the demo
is deterministic, so your numbers will match; the per-axis "Reading" column is
trimmed here for width):

```text
Judge simulated:v1 — 300 pairs
Axis             Grade  Measurement
Position         D      P(first) 63.7% [60.5%, 66.8%]
Consistency      F      flip rate 39.3% [34.0%, 44.7%]
Human agreement  C      accuracy 73.0% [69.5%, 76.7%] · κ 0.33 [0.27, 0.39]
Verbosity        D      +16.3 pts [+12.0, +20.7]
Self-preference  D      +15.4 pts [+9.4, +21.8]
Calibration      B      ECE 5.2% [3.3%, 10.0%] → 5.4% [2.5%, 9.7%]
```

Those poor grades are the point — the fixture's judge was *built* to be
biased, and the numbers above are the planted biases read back off the
instruments. (They sit a bit below the raw planted tilts because pairs where
every tilt lines up clamp at probability 1; the flagship test accounts for
that exactly. The same clamping mutes the planted overconfidence to a ~5-pt
ECE — small enough that isotonic recalibration, honestly compared on the same
held-out half, doesn't reliably improve it here; the report says so rather
than flattering the method.) See the next section.

<!-- TODO(release): add a screenshot of the HTML report card here:
     docs/report-card.png, generated from `judgemetry demo`. -->

## How it validates itself

A metrics library asks for trust; a measurement instrument should ship with a
calibration certificate. Judgemetry's demo fixture is generated from a known
model of a bad judge: base agreement with humans of 0.78, then additive tilts
of +0.18 toward the first slot, +0.22 toward the longer answer, +0.15 toward
its "own" model, and confidence inflated by 25% of the remaining headroom.
Those parameters are chosen, not observed — which means every metric has a
closed-form expected value before the pipeline ever runs.

The flagship test (`tests/test_demo_recovers_planted_biases.py`) runs the full
offline pipeline and asserts each estimate lands on its closed-form expectation
within tolerance. Without clamping those expectations would be the naive sums —
slot-1 preference = 0.5 + 0.18, verbosity bias = +0.22, consistency-aware
accuracy = 0.78 — but the default plant pushes some pair configurations past
probability 1, so the test enumerates the clamped cases exactly (expectations
of 0.637 / +0.194 / 0.723, which is what the demo's measured numbers land on
within their CIs). It also asserts that recalibration strictly improves
held-out ECE on a plant with an unambiguous overconfidence signal — before and
after are compared on the *same* held-out half, so a broken (identity)
recalibrator produces exact equality and fails — and that two runs are
byte-identical. The same pipeline runs in CI on every
commit with no secrets. So the demo is simultaneously the quick start, the CI
proof, and the evidence that the instruments measure what they claim. The
derivations and their fine print are in [DESIGN.md](DESIGN.md) — including
what this does *not* prove (that real judges behave like the synthetic one).

## Live run (real judge, real labels, real money)

```bash
pip install "judgemetry[anthropic] @ git+https://github.com/sedai77/judgemetry-llm-judge-reliability"
export ANTHROPIC_API_KEY=sk-ant-...

judgemetry fetch --dest data/          # MT-Bench human judgments (CC-BY-4.0) -> data/mtbench_pairs.jsonl
judgemetry run --dataset data/mtbench_pairs.jsonl \
    --judge anthropic:claude-opus-5 \
    --limit 200 --max-cost-usd 5 --batch \
    -o judged.jsonl
judgemetry report judged.jsonl -o report.html
```

Honest notes before you run this:

- **It costs money**, and results vary by judge model, prompt distribution,
  and dataset — the demo's numbers say nothing about any real judge.
- `--batch` uses the Anthropic Message Batches API at **50% of standard token
  prices**; without it, requests go one at a time at full price.
- Every verdict is cached (SQLite) the moment it arrives. Hitting
  `--max-cost-usd` stops cleanly (exit code 3) and a rerun resumes from the
  cache, paying only for what's missing. Reruns of a finished run are free.
  Batch runs submit in chunks of 200 judgments with the cap checked between
  chunks, so a capped batch run can overshoot by at most one chunk's cost —
  and a cap crossed on the final chunk still completes and writes the output.
- The judge sees each pair twice (positions swapped) — budget 2 judgments per
  pair.
- Human labels come from
  [`lmsys/mt_bench_human_judgments`](https://huggingface.co/datasets/lmsys/mt_bench_human_judgments)
  (CC-BY-4.0), normalized to majority labels per pair.

## The metrics

One line each; exact definitions and rationale in [DESIGN.md](DESIGN.md).

| Metric | One-liner |
| --- | --- |
| Flip rate | Fraction of pairs where swapping answer order changes the verdict — pure test–retest noise. |
| Slot-1 preference | P(winner is the first-presented answer), over judgments that picked a side; 0.5 is fair, deviation is position bias with a sign. |
| Consistency-aware accuracy | Agreement with the human label scored over both orderings; an order-dependent "correct" earns half credit. |
| Cohen's kappa | Chance-corrected agreement between the human label and the judge's cross-ordering consensus verdict. |
| Verbosity bias | P(judge prefers the longer answer) minus P(humans do) — rewards for length *beyond* human taste. |
| Self-preference | Same subtraction, on pairs where exactly one answer comes from the judge's own model family. |
| ECE / Brier | Confidence vs. empirical correctness: 10-bin expected calibration error, plus Brier as the proper-scoring cross-check. |
| Recalibration | Isotonic (PAV) map from confidence to correctness, fitted on an even/odd pair split; before and after are both evaluated on the held-out half so the comparison isolates recalibration. |

## Limitations

- **Pairwise-preference judges only.** No absolute scoring, rubric grading, or
  ranking of 3+ candidates (yet — see roadmap).
- **Single-turn comparisons.** Multi-turn conversations are out of scope.
- **One live backend so far** (Anthropic). The `Judge` protocol is a tiny
  surface — a `name` plus one `judge()` method, with an optional `judge_batch`
  for batch-capable backends — designed for third-party backends; see
  `judgemetry/judges.py`.
- **The demo data is synthetic** and is labeled as such in the report. It
  certifies the instruments, not any real judge.
- **Proxies with blind spots:** verbosity uses character length; self-
  preference uses model-name matching. [DESIGN.md](DESIGN.md) §6 lists the
  full threats-to-validity inventory.

## Related work

- [**judgecal**](https://pypi.org/project/judgecal/) — calibration tooling
  aimed at reward models and local judges in offline GPU batch settings; if
  your judge is a reward model running on your own hardware, it is the better
  fit. Judgemetry sits in a different lane: hosted-API judges scored against
  human preference labels, a bias battery (position/verbosity/self) beyond
  calibration alone, and a shareable report card.
- **MT-Bench / Chatbot Arena** (Zheng et al., 2023, *Judging LLM-as-a-Judge
  with MT-Bench and Chatbot Arena*) — the source of our human labels and the
  paper that mainstreamed position and verbosity bias in LLM judges.
- **Position-bias literature** — e.g. Wang et al., 2023, *Large Language
  Models are not Fair Evaluators*, on order sensitivity in pairwise judging;
  our two-pass swapped design is the standard mitigation turned into a
  measurement.
- **Format-restriction literature** (*Let Me Speak Freely?*-adjacent) —
  constrained/structured output can itself shift model behavior; relevant
  because Judgemetry elicits verdicts as structured JSON, so measured biases
  are properties of judge *plus* elicitation protocol.

## Roadmap

- OpenAI and local (OpenAI-compatible / llama.cpp server) judge backends.
- Rubric-decomposition scoring (judge sub-criteria, aggregate transparently).
- Cost-vs-reliability frontier: grade several judge models on the same pairs,
  plot dollars against kappa.
- CI-integration mode: a machine-readable gate ("fail the build if flip rate
  CI exceeds X") for eval pipelines.

## FAQ

**Why do judges flip when I swap answer order?**
Autoregressive scoring is not symmetric in its inputs: the first answer
conditions how the second is read, and preference for a slot (either slot)
shows up across models and prompts. It is the best-replicated LLM-judge
failure mode, which is why Judgemetry judges every pair twice by construction
and reports the flip rate first.

**Why isotonic regression and not Platt scaling?**
Platt assumes the miscalibration is sigmoid-shaped. LLM confidences cluster
on round numbers and saturate near the top — isotonic assumes only "higher
confidence should not mean lower accuracy", which is the weakest assumption
under which recalibration means anything. Trade-off and split protocol in
[DESIGN.md](DESIGN.md) §3.

**How many pairs do I need?**
Let the confidence intervals tell you: every estimate ships with a bootstrap
95% CI, and the CI width *is* the answer for your dataset and judge. As a
rough prior, rate-like metrics need a few hundred pairs for ±0.05-ish
intervals; small subgroups (self-preference especially) stay wide much
longer. Start with `--limit 200 --max-cost-usd 5`, look at the intervals,
and buy more pairs only if the question you care about is still ambiguous.

**Can I use my own dataset?**
Yes. `judgemetry run --dataset yours.jsonl` accepts JSONL where each line is:

```json
{"schema": "judgemetry/pair@1", "pair_id": "q1-m1-m2", "question": "...",
 "answer_a": "...", "answer_b": "...", "model_a": "gpt-x", "model_b": "claude-y",
 "human_winner": "A"}
```

`human_winner` is `"A"`, `"B"`, or `"tie"` for the answers **as listed**
(Judgemetry handles the position swapping itself — never pre-swap).
`model_a`/`model_b` may be `""` if unknown; self-preference is then skipped.

**Can I score verdicts from a judge Judgemetry can't call?**
Yes. The CLI only drives Anthropic judges (`--judge anthropic[:MODEL]`), but
`judgemetry report` accepts any judged-pair JSONL, so run your own judge —
each pair twice, positions swapped — and write one `judgemetry/judged@1`
object per line:

```json
{"schema": "judgemetry/judged@1",
 "record": {"pair_id": "q1-m1-m2", "question": "...", "answer_a": "...",
            "answer_b": "...", "model_a": "gpt-x", "model_b": "claude-y",
            "human_winner": "A"},
 "forward":  {"winner": "A", "confidence": 0.8, "raw": "", "refused": false},
 "backward": {"winner": "A", "confidence": 0.7, "raw": "", "refused": false},
 "judge": "my-judge:v1"}
```

`forward` is the verdict with the answers presented as `(answer_a, answer_b)`,
`backward` with them swapped — but the backward `winner` is expressed in the
**original** frame: `"A"` always means `answer_a` won, in both verdicts. A
refused judgment must be recorded as `{"winner": "tie", "confidence": 0.5,
"refused": true}`. Then: `judgemetry report judged.jsonl -o report.html`.
(Set `self_model` via the library — `metrics.compute_report_metrics(judged,
self_model="...")` — if you want self-preference for a non-Anthropic judge.)
Exact field contracts: [docs/SPEC.md](docs/SPEC.md) and `judgemetry/records.py`.

**Does the report need an internet connection to view?**
No. One HTML file, inline CSS, inline SVG, no external requests, dark/light
via `prefers-color-scheme`. Email it, attach it to a PR, archive it.

## Contributing

See [CONTRIBUTING.md](CONTRIBUTING.md) — dev setup is `uv` plus nothing, and
[docs/SPEC.md](docs/SPEC.md) is the authoritative internal contract. Security
policy in [SECURITY.md](SECURITY.md).

MIT © 2026 Judgemetry contributors — see [LICENSE](LICENSE).
