Paste model predictions with outcomes and see whether the stated
probabilities can be trusted — reliability diagram, ECE, Brier score, and
a significance test, computed entirely in your browser. Same formulas as
the calikit Python
package: pip install calikit
–
expected calibration error
n items
–
Brier score
–
log loss
–
lower is better
AUC
–
discrimination
MCE
–
worst bin gap
Spiegelhalter Z
–
Dots: mean confidence vs observed frequency per bin (hover for counts).
Dashed line: perfect calibration. Bars: items per bin.Bins table