Learn by doing

Ten quiz questions, two pretend judges, and every judgekeeper command. Each step shows the command, what it prints and what that means. It takes about fifteen minutes, needs no API key and costs nothing.

Every output on this page is real. An automatic test runs this whole tutorial on each change to judgekeeper and checks that the outputs here still match.

Before you start

  1. Install judgekeeper

    pip install judgekeeper

    Other ways to install are on the home page.

  2. Download the files into one new folder

    • items.csv: ten questions, each with an answer. Three answers are wrong: q4 says Dickens wrote Romeo and Juliet, q6 says Saturn is the largest planet, q8 says plants absorb oxygen.
    • results.csv: a pretend judge's verdicts on those answers, three times over, next to a person's labels.
    • labels-filled.csv: the ten answers, already labeled by a person.
    • lazy_judge.py and careful_judge.py: two pretend judges, written as short Python functions.
    • A folder named promptfoo with results.json and labels.csv, for the last step.

    Or clone the repository: the same files are in website/tutorial/.

  3. Open a terminal in that folder

    Every command below runs from there.

1. The demo

Start with the built-in demo. It checks a recorded judge on 100 made-up examples.

judgekeeper demo

It prints:

judgekeeper demo: a recorded judge on 100 synthetic single-output items, 3 runs.
TPR 0.97 (95% CI 0.88 to 0.99), TNR 0.86 (95% CI 0.73 to 0.94), kappa 0.84
Usable with care: TPR 0.97, TNR 0.86, kappa 0.84 against human labels. Read the flags below before relying on it.
open judgekeeper-demo/report.html (data in judgekeeper-demo/report.json)

What it means. TPR (true positive rate) is how often the judge passed what people passed. TNR (true negative rate) is how often it failed what people failed. Kappa is overall agreement after taking away lucky guessing. Here TNR is 0.86: the judge let some bad answers through, so the verdict is usable with care. The numbers in brackets are likely ranges. Open judgekeeper-demo/report.html in your browser to see the full report.

2. Check a spreadsheet

results.csv holds a judge's verdicts and a person's labels side by side, one row per verdict. The judge failed only q4. The person failed q4, q6 and q8.

judgekeeper check results.csv --judge verdict --human label --out reports/csv

It prints:

Not trustworthy as a gate: TPR 1.00, TNR 0.33, kappa 0.41 against human labels.
  - TNR is 0.33 (below 0.80): not trustworthy as a gate.
  - Kappa vs human labels is 0.41 (below 0.6): not trustworthy as a gate.
  - judge identity incomplete: provider, model, snapshot, endpoint, prompt_hash, rubric_version, temperature
  - Only 10 labeled items: error bars are wide; aim for about 100.
wrote reports/csv/report.json and reports/csv/report.html

What it means. TPR is 1.00: the judge passed every good answer. But TNR is 0.33: it caught only one of the three wrong answers. This is the most common way a judge fails: it passes too much. Raw agreement would hide it, because most answers are good. The flag "judge identity incomplete" means the spreadsheet does not say which model judged, so judgekeeper cannot tell later if the judge changed.

3. Make a labeling sheet

To check a judge you need human labels: a person's pass or fail on each answer. Start a sheet you can fill in Excel or Google Sheets:

judgekeeper template items.csv -o labels.csv

It prints:

wrote labels.csv (10 rows): fill in human_label (pass or fail), then run `judgekeeper import-labels labels.csv -o anchors.jsonl`

4. Or label in your browser

Instead of a spreadsheet, judgekeeper can open a small labeling page on your own computer. Nothing leaves your machine.

judgekeeper label items.csv --out labels.csv

A page opens in your browser and shows one answer at a time. Press 1 for pass and 2 for fail. It saves after every key, so you can stop and come back. Press Ctrl-C in the terminal when you are done.

5. Freeze the labels

This step uses labels-filled.csv, which is already labeled. If you labeled your own sheet, use labels.csv instead.

judgekeeper import-labels labels-filled.csv -o anchors.jsonl

It prints:

wrote anchors.jsonl: 10 labeled items (fail 3, pass 7), frozen (sha256 2e75fe64157b...)
  warning: Only 10 labeled items: error bars are wide; aim for about 100.

What it means. The labeled answers are now an anchor set: a fixed answer key. judgekeeper stored a hash, a fingerprint of the file. If anyone changes a label, every later command notices and stops. The warning is right: ten labels are too few for real use. Aim for about 100.

6. Run two judges, three times each

Open lazy_judge.py and careful_judge.py; each is a few lines long. The lazy judge passes any answer that sounds confident. The careful judge knows the facts, but is unsure about one question and sometimes changes its mind, as real AI judges do.

judgekeeper judge anchors.jsonl --callable lazy_judge:judge --runs 3 --out runs/lazy

It prints:

30 judge calls (10 items x 3 runs)
wrote runs/lazy/run-01.jsonl (10 judgments)
wrote runs/lazy/run-02.jsonl (10 judgments)
wrote runs/lazy/run-03.jsonl (10 judgments)
judgekeeper judge anchors.jsonl --callable careful_judge:judge --runs 3 --out runs/careful

It prints:

30 judge calls (10 items x 3 runs)
wrote runs/careful/run-01.jsonl (10 judgments)
wrote runs/careful/run-02.jsonl (10 judgments)
wrote runs/careful/run-03.jsonl (10 judgments)

What it means. --callable lazy_judge:judge means "the function judge in the file lazy_judge.py". Each run is one full pass over the ten answers. A judge that calls a real model works the same way.

7. Compare each judge with the person

judgekeeper validate anchors.jsonl runs/lazy --out reports/lazy

It prints:

Not trustworthy as a gate: TPR 0.86, TNR 0.67, kappa 0.52 against human labels.
  - TPR is 0.86 (between 0.80 and 0.90): usable with care.
  - TNR is 0.67 (below 0.80): not trustworthy as a gate.
  - Kappa vs human labels is 0.52 (below 0.6): not trustworthy as a gate.
  - judge identity incomplete: provider, model, snapshot, endpoint, prompt_hash, rubric_version, temperature
  - Only 10 labeled items: error bars are wide; aim for about 100.
wrote reports/lazy/report.json and reports/lazy/report.html
judgekeeper validate anchors.jsonl runs/careful --out reports/careful

It prints:

Usable as a gate: TPR 0.95, TNR 1.00, kappa 0.93 against human labels. Read the flags below before relying on it.
  - judge identity incomplete: provider, model, snapshot, endpoint, prompt_hash, rubric_version, temperature
  - Only 10 labeled items: error bars are wide; aim for about 100.
wrote reports/careful/report.json and reports/careful/report.html

What it means. The lazy judge passed the wrong answer about Saturn, because it sounds sure of itself. Its TNR is 0.67 and kappa 0.52: not trustworthy. The careful judge caught every wrong answer (TNR 1.00), so it is usable as a gate. Open reports/careful/report.html and find the noise section: the careful judge changed its verdict on 10% of the answers, the continents question (q7), between runs. That is why judgekeeper runs a judge more than once.

8. Gate on the result

A gate is an automatic check that can stop a release. In CI it fails the build when the judge is not good enough, using the command's exit code: 0 means pass, anything else stops the build.

judgekeeper gate reports/lazy/report.json --out gates/lazy

It prints, and exits with code 1:

FAIL: Kappa mean 0.52 is below the 0.60 threshold; TNR mean 0.67 is below the 0.80 threshold.
  no baseline: absolute thresholds only
wrote gates/lazy/gate.json and gates/lazy/gate.md (exit 1)
judgekeeper gate reports/careful/report.json --out gates/careful

It prints, and exits with code 0:

PASS: The judge clears every absolute threshold; no baseline to compare against.
  no baseline: absolute thresholds only
wrote gates/careful/gate.json and gates/careful/gate.md (exit 0)

What it means. The lazy judge would stop the build; the careful one lets it through. With a saved baseline (an earlier report you trust, set with judgekeeper baseline set reports/careful/report.json), the gate also fails when the judge gets worse than before by more than its own noise.

9. Switch judges with migrate

Pretend you are replacing the lazy judge with the careful one. migrate compares them on the same frozen answers:

judgekeeper migrate anchors.jsonl runs/lazy runs/careful --out migration

It prints:

BETTER: The new judge agrees with humans more than the old one (kappa 0.52 → 0.93, +0.40, outside the ±0.22 noise band). Switch and rebase: scores will move because the judge improved, not because your system changed, so do not compare scores across the switch.
wrote migration/migration.json and migration/migration.html

What it means. BETTER: the new judge agrees with the person more, by more than the noise between runs. The noise band is how much kappa moves on its own between identical runs; a change inside it proves nothing. Open migration/migration.html to see each answer whose verdict changed, and whether the change fixed or broke it.

10. Read results from an eval framework

promptfoo is a popular tool for testing AI apps. The promptfoo folder holds a results file that promptfoo itself wrote, for six different questions graded three times, and a person's labels for them.

judgekeeper import promptfoo promptfoo/results.json --id-var qid --labels promptfoo/labels.csv --out reports/promptfoo

It prints:

Not trustworthy as a gate: TPR 0.78, TNR 0.67, kappa 0.44 against human labels.
  - TPR is 0.78 (below 0.80): not trustworthy as a gate.
  - TNR is 0.67 (below 0.80): not trustworthy as a gate.
  - Kappa vs human labels is 0.44 (below 0.6): not trustworthy as a gate.
  - 16.7% of items changed verdict between identical runs (above 10%): noisy, use majority of runs.
  - judge identity incomplete: snapshot, endpoint, rubric_version, temperature
  - Only 6 labeled items: error bars are wide; aim for about 100.
wrote reports/promptfoo/report.json and reports/promptfoo/report.html

What it means. judgekeeper read promptfoo's own file; nothing had to be converted. This grader is weak (kappa 0.44) and noisy: it changed its mind between runs. --id-var qid tells judgekeeper which promptfoo variable names each question. DeepEval, Inspect AI, MLflow and Langfuse work the same way: see Works with what you already use.

Where next