Metadata-Version: 2.5
Name: judgekeeper
Version: 0.1.0
Summary: Validate, monitor and migrate your LLM-as-judge against a frozen human reference set.
Project-URL: Homepage, https://github.com/judgekeeper/judgekeeper
Project-URL: Changelog, https://github.com/judgekeeper/judgekeeper/blob/main/CHANGELOG.md
Author: Sathvik Thota
License-Expression: MIT
License-File: LICENSE
Keywords: calibration,drift,evaluation,llm,llm-as-judge
Classifier: Development Status :: 3 - Alpha
Classifier: Framework :: Pytest
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.11
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.40; extra == 'anthropic'
Provides-Extra: dev
Requires-Dist: build>=1; extra == 'dev'
Requires-Dist: hatchling>=1.26; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: pyyaml>=6; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Provides-Extra: inspect
Requires-Dist: inspect-ai>=0.3; extra == 'inspect'
Provides-Extra: mlflow
Requires-Dist: mlflow>=3.1; extra == 'mlflow'
Provides-Extra: openai
Requires-Dist: openai>=1.0; extra == 'openai'
Description-Content-Type: text/markdown

# judgekeeper

**Checks whether your LLM-as-judge agrees with human labels, and keeps checking:** TPR, TNR and kappa against a frozen human-labeled set, run-to-run noise, position bias, drift and a CI gate.

```
pip install judgekeeper
# before the first PyPI release: pip install "judgekeeper @ git+https://github.com/judgekeeper/judgekeeper"
judgekeeper demo
```

`demo` validates a recorded judge on 100 bundled synthetic items: no API key, no network, a few seconds. It prints TPR, TNR and kappa and writes `judgekeeper-demo/report.html`. Without installing: `uvx judgekeeper demo` (before the PyPI release: `uvx --from git+https://github.com/judgekeeper/judgekeeper judgekeeper demo`).

See a real report without installing anything: [LLMBar judged by Claude Haiku 4.5](docs/examples/llmbar-haiku/report.html), also published at https://judgekeeper.github.io/judgekeeper/examples/llmbar-haiku/ (the pattern is `https://<owner>.github.io/judgekeeper/examples/llmbar-haiku/`).

New to judge evaluation? The website in [`website/`](website/) explains it step by step, with a setup guide for API keys and a hands-on tutorial. Preview it locally: `python3 -m http.server --directory website 8000`, then open http://localhost:8000.

Already have judge verdicts and human labels in a CSV? One command:

```
judgekeeper check results.csv --judge verdict --human label --out reports/my-judge/
```

One row per judgment. `id`, `run`, `input`, `output` and `reason` columns are used when present. Verdicts can be `pass`/`fail`, `true`/`false`, `yes`/`no`, `correct`/`incorrect`, `1`/`0` or text starting with `PASS`/`FAIL`; scores need a rule (`--pass-if "score>=0.5"`); other spellings need `--label-map "good=pass,bad=fail"`. judgekeeper never guesses.

- No labels yet? `judgekeeper label items.jsonl` opens a local labeling page (keys 1/2, defer, undo, notes) that writes `labels.csv` as you go. [Details](docs/reference.md#label-a-local-labeling-page).
- Gate from pytest: `pytest --judgekeeper-report reports/my-judge/report.json`, a `judgekeeper_gate` fixture or a `@pytest.mark.judgekeeper` marker. [Details](docs/reference.md#pytest-plugin).
- Let a coding agent wire it in: [`skills/judgekeeper/SKILL.md`](skills/judgekeeper/SKILL.md) tells Claude Code and other agents how, step by step.

> Status: alpha (0.1.0). The validation report, CI gate and GitHub Action, pytest plugin, labeling page, judge migration, drift attribution and the promptfoo, DeepEval, Inspect AI, MLflow and Langfuse readers work. [`docs/reference.md`](docs/reference.md) has every command, flag, exit code and config key; [`CHANGELOG.md`](CHANGELOG.md) what is in 0.1.0.

## Why

Teams grade AI outputs with another model, the "judge". Judges disagree with humans more than raw agreement suggests, flip verdicts between identical runs, and change silently when the provider updates the model. judgekeeper measures a judge against a frozen set of human labels, reports TPR and TNR (with 95% intervals) and kappa rather than raw agreement, and tells you, later, whether the judge is still right.

The report leads with TPR (how often the judge passes what humans passed) and TNR (how often it fails what humans failed). Either below 0.80 is "not trustworthy as a gate"; 0.80 to 0.90 is "usable with care". It also shows kappa, a confusion matrix, every disagreement with the judge's rationale, the noise floor across repeated runs ("unknown" with one run, never zero), AB/BA position bias for pairwise items, per-slice numbers, and warnings when there are fewer than 60 labels or the classes split worse than 80/20.

## I have a spreadsheet

Judge verdicts and human labels already in one table: `judgekeeper check` above. Give it a `run` column (or several rows per id with different runs) to measure run-to-run noise; with one run the noise floor is reported as unknown and `gate` returns `FLAKY`. Without an `id` column, ids are a hash of input and output and the report says so. In a notebook:

```python
import judgekeeper
report = judgekeeper.check_table(df, judge="verdict", human="label")   # a DataFrame, a list of dicts or a path
```

No labels yet? Label in the local page or in Excel or Google Sheets, then freeze the labels as an anchor set:

```
judgekeeper label items.jsonl --out labels.csv        # or: judgekeeper template items.jsonl -o labels.csv
judgekeeper import-labels labels.csv -o anchors.jsonl # checks every label, skips and lists unlabeled rows, freezes
```

## I have a judge function

Run it 3 times over the anchor set (pairwise items in both AB and BA order) and get a report:

```python
import judgekeeper
def my_judge(item):          # item: id, input, output (never the human label)
    return call_my_model(item) # bool, "PASS: ...", 0.8, (verdict, reason) or {"verdict": ..., "reason": ...}
report = judgekeeper.check_judge(my_judge, "anchors.jsonl", runs=3, fingerprint={"model": "my-model"})
```

Async functions work too. A judge that raises or returns nothing is recorded as an error, excluded from the metrics and counted, never scored as a fail. From the command line, a Python function or a program in any language (one item as JSON on stdin, a verdict on stdout):

```
judgekeeper judge anchors.jsonl --callable mypkg.judges:my_judge --runs 3 --out runs/mine/
judgekeeper judge anchors.jsonl --exec "node judge.js" --runs 3 --out runs/mine/
judgekeeper validate anchors.jsonl runs/mine/ --out reports/mine/
```

judgekeeper prints the number of judge calls first and asks for `--yes` above 1,000. The `--exec` contract, with 10-line Node and Python judges, is in [`docs/reference.md`](docs/reference.md). Built-in Anthropic and OpenAI-compatible runners (`--runner anthropic|openai`, any OpenAI-compatible endpoint via `--base-url`) are documented there too. Keys stay in environment variables; nothing judgekeeper writes or prints contains one.

## I use a framework

Point judgekeeper at the files your eval tool already writes and add human labels. Your eval code does not change.

promptfoo (human labels from web-UI ratings, or `--labels`; repeats from `--repeat 3`):

```
promptfoo eval -o results.json --repeat 3
judgekeeper import promptfoo results.json --metric helpfulness --out reports/helpfulness/
```

DeepEval (one `test_run_*.json` per run in a results folder; labels from `--labels`):

```
export DEEPEVAL_RESULTS_FOLDER=deepeval-results   # then run your DeepEval tests 3 times
judgekeeper import deepeval deepeval-results/ --metric "Correctness [GEval]" --labels labels.csv --out reports/correctness/
```

Inspect AI (epochs are runs; labels from score edits or `--labels`; `.eval` logs need the `judgekeeper[inspect]` extra):

```
inspect eval task.py --epochs 3 --log-format json
judgekeeper import inspect logs/ --metric model_graded_qa --labels labels.csv --out reports/qa/
```

MLflow (judge and human assessments on your traces; each evaluation run is a run; needs the `judgekeeper[mlflow]` extra):

```
judgekeeper import mlflow --experiment my-app-eval --metric correctness --out reports/correctness/
```

Langfuse (judge and human scores from the public API; keys from `LANGFUSE_PUBLIC_KEY` and `LANGFUSE_SECRET_KEY`):

```
judgekeeper import langfuse --judge-score helpfulness --human-score helpfulness_human --from 2026-09-01 --pass-if "score>=0.5" --out reports/helpfulness/
```

Each writes the same `report.json` and `report.html` as `check`. With `--anchors-out anchors.jsonl`, the MLflow and Langfuse imports also freeze the labeled items as an anchor set, so you can re-judge them with `judgekeeper judge` for a real noise floor. How to get each file, how labels get in and each tool's traps: [`docs/integrations/`](docs/integrations/). Any other tool can export ScoreRecords (`target_id, name, annotator_kind, label, score, explanation, run, input, output, evaluator, created_at`) and use `judgekeeper import records`, with `--map` for renamed columns; `judgekeeper export records` writes judgekeeper's runs in that format.

## Keep checking

```
judgekeeper baseline set reports/my-judge/report.json   # commit .judgekeeper/baseline.json
judgekeeper gate reports/my-judge/report.json           # PASS, FAIL, FLAKY, JUDGE_CHANGED, ANCHORS_CHANGED
judgekeeper migrate anchors.jsonl runs/old/ runs/new/ --out reports/migration/
judgekeeper attribute reports/now/report.json --app-score-before X --app-score-after Y
```

`gate` never fails on a change inside the noise band. In CI, the GitHub Action (`action.yml`) runs judge, validate and gate, and the pytest plugin gates a report your pipeline already wrote. A judge field that is unknown on either side (imported data rarely records temperature or snapshot) is a warning, not a block, unless you pass `--require-fingerprint`. `migrate` compares an old and a new judge item by item; `attribute` says whether a score moved because your system changed or the judge did. Details: [`docs/reference.md`](docs/reference.md).

## First results

The judge is strong on ordinary items and weak on adversarial ones: `claude-haiku-4-5-20251001` reaches kappa 0.94 against human labels on LLMBar's Natural slice but 0.56 on Adversarial/GPTOut ([report](docs/examples/llmbar-haiku/report.html), data in [`docs/examples/llmbar-haiku/`](docs/examples/llmbar-haiku/)).

The run: LLMBar (Natural and Adversarial subsets, 419 human-labeled pairwise items) judged by `claude-haiku-4-5-20251001` at temperature 0, 3 runs, AB and BA. `scripts/llmbar_haiku_demo.sh` runs it end to end and writes the report to `docs/examples/llmbar-haiku/`.

Run on 2026-10-01 against `api.anthropic.com`. Every number below is quoted from `docs/examples/llmbar-haiku/report.json` (human-readable version: `report.html` in the same folder). TPR and TNR treat human label `A` as the positive class.

Verdict: **Usable as a gate: kappa 0.83, TPR 0.95, TNR 0.88 against human labels.**

| Run | Kappa | TPR | TNR |
|---|---|---|---|
| 1 | 0.82 | 0.95 | 0.87 |
| 2 | 0.85 | 0.96 | 0.89 |
| 3 | 0.82 | 0.95 | 0.87 |
| Mean | 0.83 | 0.95 | 0.88 |

By slice (mean over runs):

| Slice | Items | Kappa | TPR | TNR |
|---|---|---|---|---|
| Natural | 100 | 0.94 | 0.98 | 0.97 |
| Adversarial/GPTInst | 92 | 0.92 | 1.00 | 0.92 |
| Adversarial/Neighbor | 134 | 0.84 | 0.96 | 0.88 |
| Adversarial/Manual | 46 | 0.64 | 0.91 | 0.74 |
| Adversarial/GPTOut | 47 | 0.56 | 0.86 | 0.71 |

- Noise floor: mean pairwise kappa between runs 0.95. 3.6% of items changed verdict in at least one run.
- Position bias: 9.9% of judgments changed when the two outputs were swapped (AB vs BA). Kappa on the BA order alone was 0.80.
- The judge's TNR is lower than its TPR. Its errors cluster in the hardest adversarial slices (GPTOut, Manual), where kappa drops to 0.56 and 0.64.

## Why this and not X

- Vendor "align evals" features (LangSmith, Arize, MLflow, Ragas, Confident AI) do one-off calibration. judgekeeper covers what happens after: time, change, and gating.
- RAND's Judge Reliability Harness generates bias probes but has no human-kappa, no time series, no CI, and only supports OpenAI.
- Several zero-star scripts from Aug to Sep 2026 sketch anchor-set drift checks. judgekeeper aims to be the maintained, multi-provider, framework-integrated version with published numbers.

Full landscape and sources: `docs/research/00-judge-reliability-synthesis.md`.

## Roadmap

1. Core validation report on a public human-labeled dataset (done; LLMBar above).
2. GitHub Action and noise-floor gate (done).
3. Judge migration and drift attribution (built; the live migration demo is pending a key).
4. Use without a framework: `check`, `check_table`, `check_judge`, `--callable`, `--exec`, `demo`, spreadsheet labels (done).
5. Readers for promptfoo, DeepEval and Inspect AI files, and the ScoreRecord import/export format (done). MLflow and Langfuse readers, with re-judging of imported labels (done).
6. Launch readiness: labeling page, pytest plugin, agent skill, release and Pages workflows (done; publishing is the maintainer's manual step, [`docs/RELEASING.md`](docs/RELEASING.md)).
7. Demo agent (Claude Agent SDK), clarification (ask-vs-act) judge, MCP server.

## License

MIT.
