This page is generated from docs/reference.md by scripts/render_reference.py. New to judgekeeper? Start with the home page or the tutorial.

judgekeeper reference

Every command, file format, flag, exit code and config key. The README has the short version.

Anchor sets, judging and the report

judgekeeper freeze anchors.jsonl          # writes anchors.manifest.json (counts + sha256)
judgekeeper judge anchors.jsonl --runner anthropic --model claude-haiku-4-5-20251001 \
  --prompt prompts/pairwise.md --runs 3 --out runs/my-judge/
judgekeeper validate anchors.jsonl runs/my-judge/ --out reports/my-judge/

An anchor set is JSONL with id, input and human_label on every line. Pairwise items add output_a and output_b with labels A/B; single-output items add output with labels pass/fail. slice and notes are optional. After freeze, every command checks the anchor set against its manifest and exits with code 3 if it changed.

judge writes one run-NN.jsonl per run. Line 1 is a header with the judge fingerprint and a source object ({kind: judgekeeper | table | callable | exec | promptfoo | deepeval | inspect | records | mlflow | langfuse, file, metric}, plus version, notes and warnings for imports; imported and custom judges also record the --pass-if rule and label map used). Every judgment line carries the judge fingerprint (provider, model, served snapshot, endpoint, prompt hash, rubric version, temperature, timestamp). Pairwise items are judged in both AB and BA order.

validate writes report.json and a self-contained report.html. The headline is TPR, TNR (each with a 95% Wilson interval) and Cohen's kappa, mean over runs. Kappa is chance-corrected agreement with the human labels; TPR and TNR are the numbers to act on. The interval's n is the number of human-positive (or negative) items, not items × runs, because runs re-judge the same items. The report also has: the fingerprint and source, per-run numbers, a confusion matrix on majority verdicts, the noise floor across runs (per-item flip rate, run-vs-run kappa, majority verdict; with one run it reads "unknown: one run supplied", never zero), AB/BA position bias, a per-slice table, and every disagreement with the judge's rationale.

The verdict:

conditionsays
TPR or TNR below 0.80, or kappa below 0.6"not trustworthy as a gate"
TPR and TNR between 0.80 and 0.90 (and kappa at least 0.6)"usable with care"
otherwise"usable as a gate"

Flags: position bias above 0.10, more than 10% of items flipping between runs ("noisy, use majority of runs"), one run only (noise floor unknown), judge errors above 2% of judgments, an incomplete judge fingerprint ("judge identity incomplete: <fields>"), fewer than 60 labeled items ("error bars are wide; aim for about 100") and a class split more lopsided than 80/20. report.json keys read by gate are stable; session 4 added source, normaliser, errors, label_quality, notes, fingerprint_unknown, verdict.level, headline.tpr_ci/tnr_ci and noise_floor.status.

check: a table in, a report out

judgekeeper check results.csv --judge verdict --human label [--id id] [--run run] \
  [--input input --output output --reason reason] [--pass-if RULE] [--label-map MAP] --out reports/x/

In Python (pandas is optional; anything with to_dict(orient="records") works):

import judgekeeper
report = judgekeeper.check_table(df_or_rows_or_path, judge="verdict", human="label",
                                 pass_if=None, label_map=None, fingerprint={"model": "my-judge"},
                                 out="reports/x")   # out=None writes to a temporary directory

How verdicts are read

One normaliser serves check, check_judge, --callable, --exec and import-labels. It turns a raw judge output into pass/fail (or A/B) or error, plus a rationale:

A value that looks like a verdict but is not in the map (good), or a number with no rule, is a usage error that lists every such value seen. judgekeeper never guesses. A judge that raised, returned nothing (None, an empty string, NaN) or returned something unreadable (a dict with no verdict key, a list of three) is recorded as error, never as fail: errors are excluded from every agreement metric, counted in the report (errors) and flagged above 2% of judgments. Human labels use the same label map but never --pass-if.

Bring your own judge

Python:

judgekeeper.check_judge(judge, anchors, runs=3, pass_if=None, label_map=None, fingerprint=None,
                        out=None, yes=False)   # returns the report dict

judge(item: dict) returns a bool, str, float, tuple or dict; it may be async def (this also works inside a running event loop, as in a notebook). item is the anchor item without human_label and notes. Pairwise items are judged twice, with output_a and output_b swapped the second time; the second verdict is mapped back to the original labels. anchors is a frozen anchor JSONL path, or a list of items (written under out and frozen). fingerprint takes what you know: provider, model, snapshot, endpoint, prompt (text, hashed) or prompt_hash, rubric_version, temperature; everything else is recorded as unknown. Above 1,000 judge calls it needs yes=True.

Command line:

judgekeeper judge anchors.jsonl --callable mypkg.judges:my_judge --runs 3 --out runs/x/
judgekeeper judge anchors.jsonl --exec "node judge.js" --runs 3 --out runs/x/

Both print the number of judge calls (items × runs, × 2 for pairwise) before starting and need --yes above 1,000. --pass-if and --label-map work as above; --model, --temperature and --prompt (any file; hashed, with rubric_version taken from its frontmatter if it has one) fill in the fingerprint. --callable imports from the working directory and calls the function one item at a time. --exec runs the command once per item, --workers at a time (default 4), with --timeout seconds each (default 300).

The --exec contract: one item as JSON on stdin per invocation. Stdout is either a bare verdict (pass, FAIL: wrong total, 0.83) or a JSON object the normaliser reads ({"verdict": "pass", "reason": "..."}). A non-zero exit, a timeout or empty stdout is an error judgment; the last lines of stderr go into its error text (scrubbed of keys).

Node (judge.js):

// judge.js: run with --exec "node judge.js"
let data = "";
process.stdin.on("data", (chunk) => (data += chunk));
process.stdin.on("end", async () => {
  const item = JSON.parse(data); // { id, input, output } or { id, input, output_a, output_b }
  // Call your model here. This placeholder passes any non-empty output.
  const pass = item.output.trim().length > 0;
  const reason = pass ? "output is not empty" : "empty output";
  console.log(JSON.stringify({ verdict: pass ? "pass" : "fail", reason }));
});

Python (judge.py):

# judge.py: run with --exec "python judge.py"
import json
import sys

item = json.load(sys.stdin)  # {"id", "input", "output"} or {"id", "input", "output_a", "output_b"}
# Call your model here. This placeholder passes any non-empty output.
ok = bool(item["output"].strip())
reason = "output is not empty" if ok else "empty output"
print(json.dumps({"verdict": "pass" if ok else "fail", "reason": reason}))
sys.exit(0)  # a non-zero exit is recorded as an error, not a fail

Labels from a spreadsheet

judgekeeper template items.jsonl -o labels.csv      # or items.csv
judgekeeper import-labels labels.csv -o anchors.jsonl [--label-map "good=pass,bad=fail"]

template writes id,input,output,human_label,notes (pairwise: output_a,output_b) with human_label and notes empty, ready for Excel or Google Sheets. Ids come from an id column or are derived as in check. import-labels reads the sheet back, checks every label through the normaliser (an unmapped label is a usage error listing them), skips and lists unlabeled rows, keeps non-empty notes and slice, writes the anchor file and freezes it. It prints the label-quality warnings that also appear in every report.

label: a local labeling page

judgekeeper label items.jsonl [--out labels.csv] [--port 8765] [--no-browser]   # or items.csv

Opens a page in your browser that shows one item at a time: the input and the output (pairwise: A and B side by side). Keys: 1 pass (pairwise: A, also a), 2 fail (pairwise: B, also b), d defer, u undo, n note, arrow keys to move; a button mirrors every key. A judge verdict in the items (a judge_verdict, verdict or judge column, with judge_reason, reason or rationale) sits behind a "Show judge" button, hidden by default so it does not anchor the labeler.

Every change is written to --out at once (to a temporary file, then renamed over it) as id,input,output,human_label,notes (pairwise output_a,output_b), the shape template writes, so import-labels and import --labels read it unchanged. Single items get pass/fail, pairwise items A/B; deferred items stay unlabeled (deferral is not saved to the file). Reopening with the same --out resumes where it left off; an --out with ids that are not in the items is refused. items may itself be a sheet from template, partly filled in. When every item is labeled or deferred, the page shows the counts, the split, the label-quality warnings (fewer than 60 labels, worse than 80/20) and the import-labels command to run next.

The server uses only the standard library and binds 127.0.0.1, never 0.0.0.0. The URL carries a random token that the page and every request must present; a request whose Host header is not 127.0.0.1:<port> is refused (DNS rebinding); the page loads nothing from the network (a Content-Security-Policy enforces it) and shows every string as text. Ctrl-C stops it, and it stops by itself after 2 hours without a request.

import: promptfoo, DeepEval and Inspect AI results

judgekeeper import <tool> <path>... [--metric NAME] [--labels labels.csv] [--pass-if RULE] \
  [--label-map MAP] [--runs-by-order] [--id-var NAME] [--map MAP] --out reports/x/

<tool> is promptfoo, deepeval, inspect or records. A path is a file, a directory or a quoted glob; files are read in the order given (a directory or glob in name order). The output is what check writes: anchors.jsonl with its manifest, runs/run-NN.jsonl and report.json / report.html, with source.kind set to the tool and source.version to the tool's own format version when the file states one (promptfoo results.version, Inspect version). Exit 0 on success, 2 on a usage error. Per-tool pages: promptfoo, DeepEval, Inspect AI.

How records become a report: human records become anchor labels, judge records become judgments and code records (promptfoo's contains, javascript, ...) are ignored with a note. Several files, or several run indices in one file (Inspect epochs, promptfoo repeats), become separate runs; an item a run does not judge is an error judgment in that run. With one run the noise floor is "unknown: one run supplied". An item with a verdict and no human label, or a label and no verdict, is dropped, counted in source.n_judged_unlabeled / source.n_labeled_unjudged and noted in the report. Ids that had to be derived (a hash of input and output, as in check) are noted too.

Fingerprint: each judgment keeps the judge identity its record carries (model, prompt hash, temperature, timestamp). The run header records the fields every judgment agrees on; a field they disagree on is unknown in the header, with a note. A raw prompt is hashed into prompt_hash and never written. Tool warnings (promptfoo's unrecorded default grader, DeepEval's positional names) are report flags.

In Python:

import judgekeeper
report = judgekeeper.import_results("deepeval", ["deepeval-results/"], metric="Correctness [GEval]",
                                    labels="labels.csv", pass_if=None, label_map=None,
                                    runs_by_order=False, out="reports/x")
# also id_var="qid" (promptfoo) and column_map="target_id=trace_id,..." (records)

import mlflow and import langfuse

judgekeeper import mlflow --experiment NAME_OR_ID [--run-id ID]... [--metric NAME] \
  [--tracking-uri URI] [--id-from KEY] [--temperature T] [--labels labels.csv] \
  [--pass-if RULE] [--label-map MAP] [--anchors-out anchors.jsonl] --out reports/x/

judgekeeper import langfuse --judge-score NAME --human-score NAME \
  (--from DATE [--to DATE] | --to DATE | --max-items N) [--rate PER_MINUTE] \
  [--labels labels.csv] [--pass-if RULE] [--label-map MAP] [--anchors-out anchors.jsonl] --out reports/x/

These read a platform instead of files: they take no paths, and their flags are usage errors with any other tool. The output, --labels, --pass-if, --label-map and --runs-by-order are as for the file readers above. Per-platform pages: MLflow, Langfuse.

MLflow (needs pip install "judgekeeper[mlflow]"; without it, a usage error saying so):

Langfuse (standard library only):

--anchors-out PATH (both): also write every item with a human label as an anchor set (id, input, output, human_label; input and output only when the platform returned them, else ""; no rationales, no annotator ids) and freeze it, ready for judgekeeper judge. Usage error with the file readers.

In Python:

report = judgekeeper.import_results("mlflow", metric="correctness", out="reports/x",
                                    source={"experiment": "my-app-eval", "run_ids": None,
                                            "tracking_uri": None, "id_from": None,
                                            "temperature": None},
                                    anchors_out="anchors.jsonl")
report = judgekeeper.import_results("langfuse", pass_if="score>=0.5", out="reports/y",
                                    source={"judge_score": "helpfulness",
                                            "human_score": "helpfulness_human",
                                            "from_": "2026-09-01", "to": None,
                                            "max_items": None, "rate": 30})

ScoreRecords: import records and export records

A ScoreRecord is one verdict. Field names follow OpenInference annotations, so other tools' exports map onto it by renaming:

fieldmeaning
target_idthe item. Missing: derived from input and output as in check, and the report says so
namethe metric or scorer (default judge); choose one with --metric
annotator_kindLLM (a judge verdict), HUMAN (a label) or CODE (ignored). Case-insensitive; LLM_JUDGE reads as LLM. Default LLM
labelthe verdict or label: pass/fail, true/false, a bool, or anything --label-map maps
scorea number; read as the verdict with --pass-if, or when label is empty
explanationthe judge's rationale
runrepeat index, or empty. Without one, each file is one run (see --runs-by-order)
input, outputthe item's text (or any JSON)
evaluator{provider, model, prompt, temperature, version}; version is the rubric version, prompt is hashed. Also prompt_hash, snapshot and endpoint, which export records writes
created_atwhen the verdict was made

HUMAN records label the metric they are named after, or any metric when their name is not a judge metric in the file (promptfoo's ratings are named human).

judgekeeper import records records.jsonl --out reports/x/
judgekeeper import records export.csv --map "target_id=trace_id,name=metric,label=value,explanation=comment,annotator_kind=source" --out reports/x/
judgekeeper export records reports/x/ -o records.jsonl

import records reads JSONL, CSV or TSV. --map takes field=column pairs; evaluator fields are evaluator.model=judge_model and so on (a CSV can also have evaluator.model columns, or an evaluator column holding JSON). A mapped column that does not exist is a usage error listing the columns.

export records <dir> -o records.jsonl [--anchors anchors.jsonl] writes judgekeeper's runs and anchors as ScoreRecords: one HUMAN record per anchor item and one LLM record per judgment, with the judgment's fingerprint as evaluator (the prompt as prompt_hash) and error judgments as an empty label. <dir> is a directory check or import wrote (or its report.json), or a runs directory with --anchors. import records on the result gives the same report. Single-output anchor sets only; slice and notes are not carried.

demo

judgekeeper demo [--out judgekeeper-demo] validates a recorded judge on 100 synthetic single-output items (arithmetic questions, 3 recorded runs) shipped inside the package, writes report.html and report.json, and prints the path and the TPR, TNR and kappa. No key, no network. The data is made by scripts/make_demo_data.py; see src/judgekeeper/demo_data/NOTICE.

Unknown judge fields

Every fingerprint field except created_at may be unknown: provider, model, snapshot, prompt hash, rubric version and temperature are null, and an unknown endpoint is the string "unknown" (because endpoint: null already means the provider's default endpoint). Reports show such fields as "unknown" and flag "judge identity incomplete: <fields>". Files from earlier versions still load: a missing field reads as unknown, except a missing endpoint, which reads as the provider default.

gate compares the baseline and the report field by field. A field known on both sides that differs is JUDGE_CHANGED, as always. A field unknown on either side adds a warning and does not block, unless --require-fingerprint is passed, in which case the status is JUDGE_CHANGED with the reason "cannot prove same judge". The default is to warn because imported data rarely carries temperature or snapshot, and blocking on that would make the gate unusable for it; pass --require-fingerprint once your judge records its full identity.

Your API keys

A custom endpoint: any OpenAI-compatible server (Azure OpenAI's v1 endpoint, OpenRouter, Together, a LiteLLM proxy, Ollama, vLLM) with --runner openai, or a gateway in front of Anthropic with --runner anthropic. OPENAI_BASE_URL and ANTHROPIC_BASE_URL also work; the flag wins when both are set. With --base-url and no key set, judgekeeper sends a placeholder key, since local servers need none. A URL with user:password@ in it is rejected.

judgekeeper judge anchors.jsonl --runner openai --model "$JUDGE_MODEL" \
  --base-url http://localhost:11434/v1 --prompt prompts/pairwise.md --runs 3 --out runs/local/

A custom key variable:

export OPENROUTER_API_KEY=...            # in your shell or CI secrets, never in a file
judgekeeper judge anchors.jsonl --runner openai --model "$JUDGE_MODEL" \
  --base-url https://openrouter.ai/api/v1 --api-key-env OPENROUTER_API_KEY \
  --prompt prompts/pairwise.md --runs 3 --out runs/openrouter/

The endpoint's host goes into the fingerprint as endpoint (null for the provider default): the same model name behind a different endpoint is a different judge, and gate treats a changed endpoint as JUDGE_CHANGED. AWS Bedrock and Google Vertex have no native client; put a gateway such as LiteLLM in front of them and use --base-url.

Gate CI on the judge

judgekeeper baseline set reports/my-judge/report.json   # copies to .judgekeeper/baseline.json; commit it
judgekeeper baseline show                               # fingerprint, anchors hash, kappa / TPR / TNR
judgekeeper gate reports/my-judge/report.json           # writes gate.json and gate.md next to the report

gate compares a report with fixed thresholds and, if there is one, the baseline (--baseline, default .judgekeeper/baseline.json when it exists). It decides one status, checking in this order:

  1. ANCHORS_CHANGED: the anchor set hash differs from the baseline's.
  2. JUDGE_CHANGED: provider, model, snapshot, endpoint, prompt hash, rubric version or temperature differ from the baseline. Scores are not compared across judges. --allow-judge-change turns this into a warning and gates on absolute thresholds only. A field unknown on either side is a warning, or JUDGE_CHANGED ("cannot prove same judge") with --require-fingerprint; see Unknown judge fields.
  3. FLAKY: fewer than 3 runs, so the noise floor is unknown (with one run: "noise floor unknown: one run supplied").
  4. Absolute thresholds: kappa mean >= 0.6, TPR mean >= 0.8, TNR mean >= 0.8, AB/BA disagreement <= 0.10.
  5. Against the baseline: kappa, TPR and TNR may not drop by more than the noise band, which is the larger of the baseline's and this report's run-to-run spread, and at least 0.02.
  6. A check that fails on the mean but passes on the best run, or more than 10% of items flipping between runs, is FLAKY ("use the majority of more runs") instead of FAIL.
  7. Otherwise PASS.

JUDGE_CHANGED points at migrate, below: compare the old and new judge, then --rebase to make the new one the baseline.

--flaky-as pass or --flaky-as fail maps FLAKY to exit 0 or 1; the status in gate.json and gate.md stays FLAKY. Without the flag FLAKY exits 4. gate.md is a short summary for a PR comment or $GITHUB_STEP_SUMMARY: the status, why, a metric / baseline / now / delta / noise band table and the judge fingerprint.

Thresholds live in an optional judgekeeper.toml (--config, default ./judgekeeper.toml when it exists), one table per command: [gate] here, [migrate] and [attribute] below. Unknown tables and keys are a usage error.

[gate]
kappa_min = 0.6
tpr_min = 0.8
tnr_min = 0.8
ab_ba_disagreement_max = 0.10
min_band = 0.02        # smallest noise band, for judges that never vary between runs
flip_rate_max = 0.10   # share of items that may flip between runs before a failure is FLAKY

pytest plugin

Installed with judgekeeper (pytest11 entry point judgekeeper.pytest_plugin); pip install pytest-judgekeeper installs the same thing under the name the pytest plugin list uses. It runs the gate logic on a report.json your pipeline already wrote. It never calls a judge or an API and writes nothing.

def test_judge_still_agrees_with_humans(judgekeeper_gate):
    judgekeeper_gate("reports/my-judge/report.json")   # baseline=None, config=None, allow=("PASS",), flaky_as=None

@pytest.mark.judgekeeper(report="reports/my-judge/report.json", flaky_as="pass")
def test_with_the_marker():
    ...   # runs only if the gate allows the report
pytest --judgekeeper-report reports/my-judge/report.json [--judgekeeper-baseline B] \
  [--judgekeeper-config judgekeeper.toml] [--judgekeeper-flaky-as pass|fail]

Migrate to a new judge

When a judge model is deprecated, or you want a cheaper or better one, judge the same frozen anchor set with both and compare:

judgekeeper judge anchors.jsonl --runner anthropic --model "$OLD_MODEL" --prompt prompts/pairwise.md --runs 3 --out runs/old/
judgekeeper judge anchors.jsonl --runner anthropic --model "$NEW_MODEL" --prompt prompts/pairwise.md --runs 3 --out runs/new/
judgekeeper migrate anchors.jsonl runs/old/ runs/new/ --out reports/migration/ [--rebase] [--fail-on worse|different]

Both run directories must have been judged against the same frozen anchor set (exit 3 otherwise). migrate writes migration.json and a self-contained migration.html with:

  1. Who changed: both fingerprints side by side, changed fields highlighted, served snapshots.
  2. Each judge against humans: kappa, TPR and TNR, mean and range over runs.
  3. Old judge against new judge: kappa between the two judges' majority verdicts, the share of items whose verdict changed, and for each change whether both judges were stable on it (unanimous across their own runs) or either was flipping anyway ("within noise"). Changed-and-stable items are fixed (the new judge now agrees with the human label) or broken, with the net.
  4. Per slice: old kappa, new kappa, delta, changed-and-stable count.
  5. Changed items: id, slice, human label, old and new verdict, fixed or broken, both rationales.
  6. Pass-rate bridge: a = P(new says pass | old said pass) and b = P(new says pass | old said fail) on the anchor set, with Wilson 95% intervals, overall and per slice, so that new_rate ≈ old_rate × a + (1 − old_rate) × b. It assumes your production traffic resembles the anchor set. Pairwise sets use "prefers A" for "pass".
  7. Status, with a sentence on what to do. The noise band is the larger of each judge's run-to-run kappa spread and min_band.
statusmeaningwhat to do
EQUIVALENTkappa vs humans moved within the noise band, and at most 2% of items changed while both judges were stablekeep comparing scores across the switch; rebase
BETTERkappa vs humans rose by more than the noise bandswitch and rebase; scores move because the judge improved
WORSEkappa vs humans fell by more than the noise banddo not migrate yet
DIFFERENTkappa within the band, but more than 2% of items changed while both judges were stableas good, on different items: old and new scores are not comparable item by item; rebase if you switch

migrate exits 0 when the analysis completes, whatever the status. --fail-on worse exits 1 on WORSE; --fail-on different exits 1 on WORSE or DIFFERENT. --rebase copies the new judge's report to the baseline (--baseline, default .judgekeeper/baseline.json) and keeps migration.json as .judgekeeper/migrations/<UTC date>-<old model>-to-<new model>.json, the audit trail. Commit both; after a rebase gate no longer reports JUDGE_CHANGED.

[migrate]
min_band = 0.02            # smallest kappa noise band
max_changed_share = 0.02   # changed-and-stable share above which equal kappa is DIFFERENT

scripts/migration_demo.sh runs this on LLMBar: RUNNER, OLD_MODEL and NEW_MODEL come from the environment, and it has no default model ids. Take them from the provider's current models and deprecations pages.

> TODO: the live migration demo is pending. No API key was available in the build environment, so no migration report is committed yet and no migration numbers appear here. When one runs, its report goes under docs/examples/.

Attribute a score change

Your app's eval score moved. Did your system change, or did the judge? The anchor set is frozen, so the outputs being judged are identical every time: any movement in verdicts on it comes from the judge. Re-judge the anchor set, validate, and compare with the baseline:

judgekeeper attribute reports/now/report.json [--baseline .judgekeeper/baseline.json] \
  [--app-score-before X --app-score-after Y]

It compares per-item majority verdicts between the baseline and the current report. An item counts as moved only if it was stable (unanimous across runs) in both. Judge drift is declared when more than 2% of items moved, or kappa vs humans moved outside the noise band. This works when the declared fingerprint is identical, which is what a silent provider-side update looks like, and it reports whether the served snapshot changed.

Both reports need per-item verdicts (items, written by validate from this version on); an older baseline is a usage error that tells you to regenerate it. attribute writes attribution.json and attribution.md (for $GITHUB_STEP_SUMMARY) next to the report, or under --out. The example workflow in docs/examples/workflows/judge-gate.yml runs it weekly after the gate job re-judges the anchor set.

[attribute]
min_band = 0.02          # smallest noise band, for kappa and for app scores
max_moved_share = 0.02   # share of stable items that may move before it is JUDGE_DRIFT

Exit codes

exit codemeaningcommands
0success; PASS; STABLE; migrate finishedall
1FAIL; a judge call failed; migrate --fail-on matched; a Langfuse request failedgate, judge, migrate, import langfuse
2usage error (bad arguments, report, baseline or config; unmapped verdicts or labels; duplicate ids; several metrics without --metric; more than 1,000 judge calls without --yes; a missing extra or platform key; a Langfuse import with no time window or --max-items)all
3anchor set changed: hash mismatch, ANCHORS_CHANGED, runs or reports from a different anchor setall that read anchors or reports
4FLAKY (--flaky-as maps it to 0 or 1)gate
5JUDGE_CHANGEDgate
6JUDGE_DRIFTattribute
7SYSTEM_CHANGEattribute

GitHub Action

action.yml at the repo root runs judge, validate and gate, appends gate.md to the job summary, uploads report.html, report.json and gate.json as an artifact, and exits with the gate's exit code. API keys come from the caller's env, never from inputs; api-key-env names the variable when it is not the provider's standard one. A copy-paste workflow that runs weekly and on PRs touching the judge prompt is in docs/examples/workflows/judge-gate.yml:

name: judge gate
on:
  schedule:
    - cron: "17 6 * * 1"   # weekly: catches provider-side judge drift between PRs
  pull_request:
    paths: ["prompts/judge.md", "evals/anchors.jsonl", ".judgekeeper/baseline.json", "judgekeeper.toml"]
jobs:
  gate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: judgekeeper/judgekeeper@main   # pin to a release tag or commit sha
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
        with:
          anchors: evals/anchors.jsonl
          runner: anthropic
          model: claude-haiku-4-5-20251001
          prompt: prompts/judge.md
          runs: 3
          flaky-as: pass
inputdefault
anchorsrequiredfrozen anchor set JSONL, manifest next to it
runnerrequiredanthropic, openai or replay
model, promptjudge model id and prompt (anthropic, openai)
fixturerecorded judgments JSONL for replay (no API key)
base-urlendpoint instead of the provider default (anthropic, openai)
api-key-envname of the variable holding the key, never the key itself
runs3judge runs; the gate needs at least 3
baseline.judgekeeper/baseline.json if presentbaseline report
configjudgekeeper.toml if presentgate thresholds
flaky-aspasspass, fail, or empty to keep exit code 4
out-dirjudgekeeper-outwhere runs, report and gate files go
artifact-namejudgekeeper-gateuploaded artifact name
python-version3.12Python used to run judgekeeper

Outputs: status and exit-code. This repo's own CI runs the action on replay fixtures on every push.