{% extends "base.html" %} {% block title %}Compare Runs{% endblock %} {% block content %}

📊 Compare Runs

← Back to Runs {% if run_type == "training" and not mixed_types %} {% endif %}
{% if mixed_types %} {# ADR-0001: cross-type Compare is rejected with a typed empty-state. The runs exist but are not comparable; list each run with its type so the user can re-select a same-type set. No metrics/metadata table renders here. #}
Runs are of different types and cannot be compared.

Compare scopes runs of a single type (training, etl, eda, or inference). Select runs of the same type to compare them side by side.

{% elif run_type == "training" %} {# Server-rendered comparison table (values must appear in the HTML response, not injected by client-side JS, so the page is usable without JS and testable via the response body). #}
{% for run in runs %} {% set metrics = (run.evaluation or {}).get("metrics") or {} %} {% set is_multi = run.is_multi %} {% endfor %}
Run ID Model(s) AUC F1 Precision Recall Threshold
{{ run.run_id }} {% if is_multi and run.evaluation.ranking %} Ensemble ({{ run.evaluation.ranking | length }} models) {% else %} Single Model {% endif %} {{ format_metric(metrics.get("auc"), best.auc, run.run_id) }} {{ format_metric(metrics.get("f1"), best.f1, run.run_id) }} {{ format_metric(metrics.get("precision"), best.precision, run.run_id) }} {{ format_metric(metrics.get("recall"), best.recall, run.run_id) }} {{ format_metric(metrics.get("threshold"), None, run.run_id) }}
{# Per-run ensemble ranking tables (multi-model runs only). #} {% for run in runs %} {% if run.is_multi and run.evaluation.ranking %}

Ensemble Ranking: {{ run.run_id }}

{% for model in run.evaluation.ranking %} {% set mm = model.metrics or {} %} {% endfor %}
RankModelAUCF1
{{ loop.index }} {{ model.name }} {{ format_metric(mm.get("auc"), None, run.run_id) }} {{ format_metric(mm.get("f1"), None, run.run_id) }}
{% endif %} {% endfor %} {# JSON data island for client-side CSV export (table itself is server-rendered). #} {% else %} {# ADR-0001: homogeneous non-training compare renders a metadata table. These run types (etl/eda/inference) do not carry training metrics, so we surface timestamp/duration/status/output_paths/feature_count instead. #}
Scope: comparing {{ run_type }} runs. Training metrics (AUC, F1, precision, recall) are not produced for this run type.
{% for entry in meta_runs %} {% set meta = entry.metadata or {} %} {% set outputs = meta.get("output_paths") or {} %} {% endfor %}
Run ID Timestamp Duration (s) Status Feature count Outputs
{{ entry.run_id }} {{ meta.get("timestamp") or "—" }} {{ meta.get("duration_seconds") if meta.get("duration_seconds") is not none else "—" }} {{ meta.get("status") or "—" }} {{ meta.get("feature_count") if meta.get("feature_count") is not none else "—" }} {% if outputs %}
    {% for key, path in outputs.items() %}
  • {{ key }}: {{ path }}
  • {% endfor %}
{% else %}—{% endif %}
{% endif %} {% endblock %}