Your AI judge says
96 %
confident.
Is it right 96 % of the time?
terminal
$ judge-audit run tasks.jsonl --judge jev
judge=jev n=120 accuracy=66.7% ece=0.3181
Shadow mode. Your humans already decided. The judge must match them.
docs/audit-jev-router.mdjev
Read this first
Hard tasks routed to the cheap model: 40 / 40
Median confidence on those mistakes: 0.96
= a coin glued to “easy”. Accuracy 66.7 % is the constant-classifier baseline.
docs/audit-jev-router-ablation.mdjev
Same judge. Same 120 tasks. Two descriptive sentences on the options.
37 / 40 routed right
Confidence when wrong: 0.59
Attack success: 0 / 40
The failure was the prompt. Only a calibration audit shows it.
Calibration, not accuracy.
Every number from a committed raw checkpoint.
Any judge: Jev, OpenJev, Claude, yours.
ECE · ZERO-ERROR COVERAGE · COST · P99 · DRIFT GATE
judge-audit
When it says 90 %, is it right 90 % of the time?
pip install kunko-judge-audit
kunko-ai-labs/judge-audit · Apache-2.0