Read the metric, then the score
-
Accuracy and acc_norm are different. Accuracy chooses the
continuation with the highest total score. Length-normalized accuracy
adjusts for continuation length. A percentage from one metric cannot stand
in for the other.
-
Full applies to the selected task. A full PIQA result
does not mean every benchmark was evaluated. Missing scores are not
estimated. Smoke tests are listed separately.
-
Reports describe an exact checkpoint. Use the saved step,
SHA-256, dataset revisions, sample count and harness version when
comparing results. These are recorded Track evaluations, not independent
certifications.
-
INT Index requires all five component scores. The Open
SLM formula uses chance-normalized HellaSwag, mean ARC Easy/Challenge,
PIQA and ArithMark 3, with weights 1, 1, 1 and 0.65. BananaMind 1.1 Elo
uses that benchmark's fixed item ratings and weights.
-
ArithMark 3 has an option-position limitation. In the
pinned 1,000-item release, correct numeric answers are always one of the
two middle values. INT retains the published formula's nominal 25% chance
floor. Read it alongside the individual scores.
-
Native models predict UTF-8 bytes. Track scores every
answer byte with up to 512 bytes of context. Tokenizer and context
differences matter when comparing with model-card results from other
systems.
Charts are drawn from the numeric records shown here. New completed
evaluations of public runs appear automatically. Making a run private
removes its reports from this page.