{# Quality tab (formerly "Evals"). Redesigned 2026-08-14 per founder feedback: the old 12-tile grid + hash-list table read as vaporbox. Reshaped around the single question "is my agent doing good work?" — one graded letter card, ranked failure patterns, plain-English rough runs. Inline eval builder ("Prevent this →") lives on each rough run so a real trace can become a persistent check without leaving the tab. DOM ids here are what clawmetry/static/js/app.js::loadEvalsTab reads after the redesign. The tab is still selected via data-tab="evals" from the sidebar so the URL / muscle-memory stays stable; only the surface inside changed. See clawmetry/quality.py + routes/quality.py. #}
Reading your agent's recent work.
Ranked by what it cost you.
Click any to see the trace, or turn one into a check.
A few runs picked at random each night. Was the agent right?
Already using an evaluation platform?
Point it at this endpoint to pull every run with its outcome, cost and tokens:
/api/otel/export?shape=sessions&window=7d