Annotator ability
Estimated from agreement patterns alone — no gold labels. 1.0 is the
prior for a new annotator, 0 means labels carry no information, negative means
systematically wrong. Whiskers are ±1 approximate standard error: judge nobody
on a wide whisker.
Likely codebook bugs
Items with strongly negative discrimination: your most able
annotators disagree with the crowd consensus. That pattern usually means the guideline
is wrong or ambiguous for this item, not the annotators. Fix the codebook, then revisit.
Item measurements
Most uncertain first — these are where the next annotation is worth
the most. The dark band on each confidence bar is the ±1 SE ability-sensitivity
interval. Resolved items have passed the confidence threshold and stop consuming
annotation budget under adaptive routing.
Study designer
Before you spend: Monte Carlo power analysis with statistical synthetic
annotators. How many annotators per item until the 95% interval on Krippendorff’s
α is narrow enough to defend? Accuracy is best taken from a small pilot.
Simulation is seeded and deterministic; ~a second to run.