Full UofA credibility report card mockup: identity header, completeness and sufficiency verdicts, factor grid, weakener findings, raidex furnished evidence with RAI badge and radar, and a signed-bundle verify footer

meta-llama/Llama-3.3-70B-Instruct

Meta · revision 6f6073b · assessed 2026-11-14

model-credibility v1.0.2

raidex constituents v3.5

extraction: LLM — anthropic/claude-sonnet-4-6 · all fields provenance-stamped

[1] Documentation completeness

11 / 17 factors present

NIST AI RMF 1.0

[2] Evaluation sufficiency

5 weakeners · 3 high

NIST AI 800-3 / V&V 40 pattern

[1] Documentation factors

filled = documented · hover for factor

[2] Evaluation sufficiency findings

high W-EV-UQ-01 — scores reported without uncertainty common: 92% of assessed models
high W-EV-NULL-04 — no null / chance baseline stated common: 88%
high W-EV-DET-03 — determinism floor not characterized common: 95%
medium W-EV-CAP-06 — capability confound not partialled common: 95%

Findings describe the published record, not the model. Methodology

[3] Furnished evidence (raidex)

RAI score 81.0 coverage 9/9 · v3.5

Furnished composite — assessed as evidence under [2], not a UofA verdict. raidex.ai

ConstituentReportedFurnishedΔ
StrongREJECT98.8unreported
SimpleQA86.083.2−2.8
BBQ35.3unreported
Raidex dimension radar Six responsibility dimensions with normalized scores Factuality Fairness Safety Sycophancy Ethics Security

Normalized 0–100 · v3.5

uofa verify llama-3.3-70b.bundle.jsonld