| Component | Full-suite size | Metric | Weight |
|---|---|---|---|
| Held-out LM | ~1M UTF-8 byte tokens | NLL / BPB | 40% |
| BLiMP-fast | 8,509 pairs | Logprob margin + accuracy | 35% |
| BLiMP Supplement | 5,218 pairs | Margin + accuracy | 10% |
| EWoK-fast | 1,500 items · HF access pending | Preference margin | 5% |
| ARC-Easy | 1,000 items | MC logprob | 10% |
Four components are ready. EWoK requires approved Hugging Face access; the combined FastScore remains unavailable until all five have results. BLiMP uses the BabyLM 2024 vocabulary-filtered data. Smoke runs use 10 items per component and 8,192 LM bytes.
LM: 1 − BPB / 8. Pair diagnostics: mean tanh of the logprob margin per maximum candidate byte length, in bits. ARC-Easy: mean (P(correct) − 1/K) / (1 − 1/K). Apply the weights above only when all components are present. These are internal diagnostics, not official benchmark scores; raw metrics remain visible.