Download export (JSON)
Verdicts + raw predictions. Majority vote fails on hard items; surprise doesn’t.
Where the crowd is likely wrong
Items whose surprisingly-popular label differs from the majority label: an answer was more popular than its own supporters predicted, the statistical signature of an informed minority. Review these first.
Annotator calibration
Calibration error: how far each annotator’s predicted agreement is from the agreement they actually got from peers (lower is better). SP-alignment: how often their label matches the surprisingly-popular verdict.
All items with verdicts
Every instance with at least the minimum number of predictions.