Benchmark guide · English fast ladder
Component Full-suite size Metric Weight
Held-out LM ~1M UTF-8 byte tokens NLL / BPB 40%
BLiMP-fast 8,509 pairs Logprob margin + accuracy 35%
BLiMP Supplement 5,218 pairs Margin + accuracy 10%
EWoK-fast 1,500 items · HF access pending Preference margin 5%
ARC-Easy 1,000 items MC logprob 10%

Four components are ready. EWoK requires approved Hugging Face access; the combined FastScore remains unavailable until all five have results. BLiMP uses the BabyLM 2024 vocabulary-filtered data. Smoke runs use 10 items per component and 8,192 LM bytes.

Experimental FastScore normalization

LM: 1 − BPB / 8. Pair diagnostics: mean tanh of the logprob margin per maximum candidate byte length, in bits. ARC-Easy: mean (P(correct) − 1/K) / (1 − 1/K). Apply the weights above only when all components are present. These are internal diagnostics, not official benchmark scores; raw metrics remain visible.