Audit report — jev

200 decisions · accuracy 95.5% · ECE 0.0389
cost $0.0038 · p50 0.741s · p99 3.202s

judge `jev` · model `typesafe-ai/jev` · backend `gateway` · recomputed 2026-09-19T09:09:12+00:00 (original run time not recorded) · judge-audit 0.2.0

Can I automate this?

Zero observed errors through the most confident 73.0% (146 decisions, confidence ≥ 0.92).
Retrospective on this dataset — not a production guarantee.

Reliability diagram

reliability diagram

Accuracy vs coverage

accuracy coverage curve
coverageaccuracymin confidencen
5%100.0%1.0010
15%100.0%1.0030
25%100.0%1.0050
35%100.0%1.0070
45%100.0%1.0090
55%100.0%0.99110
65%100.0%0.96130
75%99.3%0.90150
85%99.4%0.78170
95%97.9%0.56190

Calibration bins

binavg confidenceaccuracyn
0.5-0.60.53766.7%15
0.6-0.70.64462.5%8
0.7-0.80.750100.0%10
0.8-0.90.853100.0%16
0.9-1.00.98899.3%151

A perfectly honest judge sits on the diagonal: avg confidence == accuracy in every bin.