Polish models ranked by MultiBLiMP-PL margin over the 50% chance baseline ((accuracy−0.5)×2; below-chance stays negative, not clamped). English and Polish are different competencies, so Polish-first models are ranked on Polish tasks only — never on the English board above. The other Polish columns are a cross-check (pl_induction) and diagnostics (pl_lm, pl_multiblimp) — not part of the rank.
pl_lm BPB is contamination-uncontrolled: the out-of-sample text is public Polish Wikipedia, which appears in many models' pretraining, so a lower BPB can reflect memorisation rather than ability. It is shown for information only and never used for ranking.
Primary metric is contamination-reduced, not per-model controlled:
MultiBLiMP-PL uses natural text, so it sits between the contamination-immune
pl_induction (synthetic copy) and the uncontrolled
pl_lm (Wiki-BPB). We do not silently assume the primary is
contamination-free.
pl_induction cross-check (enforced): it is
contamination-immune. Rows whose rank margin and
pl_induction signal disagree in direction are flagged
⚖ — that model's ranking validity is questionable
and it is a trigger for a future composite. Convergence corroborates that the
rank reflects competence, not memorisation. (Directional flag is live now; a
magnitude threshold is deferred to the metrics owners once cross-model numbers
exist.)