Your AI judge says
96 %
confident.
Is it right 96 % of the time?
terminal
$
j
u
d
g
e
-
a
u
d
i
t
r
u
n
t
a
s
k
s
.
j
s
o
n
l
-
-
j
u
d
g
e
j
e
v
judge=
jev
n=
120
accuracy=
66.7%
ece=
0.3181
Shadow mode. Your humans already decided. The judge must match them.
docs/audit-jev-router.md
jev
Read this first
Hard tasks routed to the cheap model:
40 / 40
Median confidence on those mistakes:
0.96
= a coin glued to “easy”. Accuracy 66.7 % is the constant-classifier baseline.
docs/audit-jev-router-ablation.md
jev
Same judge. Same 120 tasks. Two descriptive sentences on the options.
37 / 40
routed right
Confidence when wrong:
0.59
Attack success:
0 / 40
The failure was the prompt. Only a calibration audit shows it.
Calibration
, not accuracy.
Every number from a committed
raw checkpoint
.
Any judge: Jev, OpenJev, Claude,
yours
.
ECE · ZERO-ERROR COVERAGE · COST · P99 · DRIFT GATE
judge
-
audit
When it says 90 %, is it right 90 % of the time?
pip install kunko-judge-audit
kunko-ai-labs/judge-audit · Apache-2.0