{{ summary.finished_count }}/{{ summary.task_count }} tasks complete
{% extends "layout.html" %} {% block title %}Evaluation {{ run.run_id }} · cua-speedrun{% endblock %} {% block content %}
Evaluation {{ run.run_id }} · {{ run.submission.track }} · created {{ run.created_at.strftime("%Y-%m-%d %H:%M") }} UTC
{{ summary.finished_count }}/{{ summary.task_count }} tasks complete
Passed and failed tasks
Executor start to terminal report.
{% elif run.started_at %}runningFinal elapsed time appears at completion.
{% elif run.stage == "cancelled" %}not startedStopped before an executor claimed it.
{% else %}queuedWaiting for an executor start timestamp.
{% endif %}{{ "%.1f" | format((run.result.get('mean_score') if run.result.get('mean_score') is not none else run.result.success_rate) * 100) }}% score · {{ "%.1f" | format(run.result.success_rate * 100) }}% exact success · open diagnosis
{% else %}pendingScoring starts only after task evidence is complete.
{% endif %}{% for key, value in summary.usage | dictsort %}{{ key }} {{ value }}{% if not loop.last %} · {% endif %}{% endfor %}
{% else %}—No template cost snapshot received yet.
{% endif %}Open a task to inspect timing, actions, screenshots, and logs.
Task instances appear here as the run fans out.
{% endif %}Trajectories, sandbox identifiers, and task logs.
Logs and screenshots appear as the executor reports task evidence. Use the event stream below while the evaluation is running.
{% endif %} {% if artifacts.sandboxes %}| Task instance | Environment sandbox | Agent sandbox |
|---|---|---|
| {{ sandbox.task }} | {{ sandbox.env or "—" }} | {{ sandbox.agent or "—" }} |
{{ "Partial metadata for this legacy run." if contract.legacy else "Settings fixed before execution." }}
{{ variable.name }}{{ "This evaluation" if variable.source == "evaluation" else "Saved" }}{{ contract.season_key or "assigned after execution" }}Raw execution status events.