Task outcomes count whether a run finished successfully. The judge score checks how the work was done and needs a configured judge. Open Quality for the evidence →
🔬 Evaluators on your agent
Open Evals →
🧐 Run health
red = error · yellow = waste · green = clean
⚖️ Did the change help?
Select a comparison. The result includes the number of sessions used.
Looking for recent changes...
⚙️ Advanced: compare two runs by id
Paste the session IDs. Green indicates improvement. Red indicates a worse result.
🪥 Error triage
Mute expected errors to exclude them from the counts.