Overall score is a weighted composite of six dimensions.
Fairness, privacy,
security, and robustness are measured live against a published prompt corpus
(bias probes, adversarial jailbreak attempts, factual-accuracy checks, PII-leakage
scans).
Compliance and authenticity are held at a neutral placeholder —
they don't have a clean behavioral signal observable from a single model's text
output, so rather than fabricate one, this leaderboard discloses that plainly instead
of hiding it in the number. Full detail in the
published methodology.
Want the exact prompts that caused a score to drop? That's the paid diagnostic —
see
GET /api/leaderboard/{model}/{provider}/diagnostic.