Audit report — jev

200 decisions · accuracy 100.0% · ECE 0.0036
cost $0.0037 · p50 0.863s · p99 7.494s

judge `jev` · model `typesafe-ai/jev` · backend `gateway` · recomputed 2026-09-19T09:09:12+00:00 (original run time not recorded) · judge-audit 0.2.0

Can I automate this?

Zero observed errors through the most confident 100.0% (200 decisions, confidence ≥ 0.89).
Retrospective on this dataset — not a production guarantee.

Reliability diagram

reliability diagram

Accuracy vs coverage

accuracy coverage curve
coverageaccuracymin confidencen
5%100.0%1.0010
15%100.0%1.0030
25%100.0%1.0050
35%100.0%1.0070
45%100.0%1.0090
55%100.0%1.00110
65%100.0%1.00130
75%100.0%1.00150
85%100.0%1.00170
95%100.0%0.98190

Calibration bins

binavg confidenceaccuracyn
0.8-0.90.890100.0%2
0.9-1.00.998100.0%198

A perfectly honest judge sits on the diagonal: avg confidence == accuracy in every bin.