{# Quality tab (formerly "Evals"). Redesigned 2026-08-14 per founder feedback: the old 12-tile grid + hash-list table read as vaporbox. Reshaped around the single question "is my agent doing good work?" — one graded letter card, ranked failure patterns, plain-English rough runs. Inline eval builder ("Prevent this →") lives on each rough run so a real trace can become a persistent check without leaving the tab. DOM ids here are what clawmetry/static/js/app.js::loadEvalsTab reads after the redesign. The tab is still selected via data-tab="evals" from the sidebar so the URL / muscle-memory stays stable; only the surface inside changed. See clawmetry/quality.py + routes/quality.py. #}
Quality this week
Loading…
{# ── The report card: giant stamped letter + one sentence + week dots ── #}

Loading grade…

Reading your agent's recent work.

{# ── Two panels: patterns + rough runs. Both read from the same fetch. ── #}

What went wrong

Ranked by what it cost you.

    The rough runs

    Click any to see the trace, or turn one into a check.

      {# ── Footer strip: judge upgrade path. Hidden once a key is set. ── #}
      {# end page-evals #}