Today Loading task outcomes...
show details
⚡ Agent efficiency playbook
Use measured evidence to get more useful work from your agents.
1. Parallelize independent work
Activity shows active sessions, throughput, and unproductive runs.
Open Activity →
2. Continue work in the same session
Context usage shows prompt reuse, compactions, and what was measured.
Open Context usage →
3. Let unattended work be watched
Guard and Sessions show what is still running, waiting, or needs review.
Open Guard →
4. Route by outcome and cost
Models, Quality, and Cost together show whether a routing change helped.
Review routing →
🎯 How independent is your agent?
Unknown
Unknown
Last 7 days
No data yet
LLM judge quality
Unknown
Task outcomes count whether a run finished successfully. The judge score checks how the work was done and needs a configured judge. Open Quality for the evidence →
🔬 Evaluators on your agent
Open Evals →
⚖️ Did the change help?
Select a comparison. The result includes the number of sessions used.
Looking for recent changes...
⚙️ Advanced: compare two runs by id
Paste the session IDs. Green indicates improvement. Red indicates a worse result.
🪥 Error triage
Mute expected errors to exclude them from the counts.
LIVE Loading...
auto-refreshes every 30s
Loading flow...
🏥 System Health
Services
Channels
Disk Usage
Cron Jobs
Sub-Agents (24h)
Heartbeat
🔍 Configuration Diagnostics
▼
Loading diagnostics...
🐝 Active Tasks
⟳ 30s
🐝
Loading tasks...
🧠 Claude Opus ...
Waiting for activity...
Unknown Unknown Unknown
Unknown
Unknown
Unknown
❤️ Is your agent alive? ...
waiting...
Last check-in: Unknown
Unknown
Unknown
Recent check-ins
Unknown