{% set sr = sweep_report %} {% set kind = sr.objective_kind %}

Cheapest verified configuration sweep · minimize {{ kind }} subject to verified pass rate ≥ {{ sr.min_pass_rate|pct }} · {{ sr.n_configs }} configurations · {{ sr.n_runs }} runs{% if sr.n_skipped %} · {{ sr.n_skipped }} skipped by budget{% endif %}

{% for note in sr.notes %}

{{ note }}

{% endfor %}
{% for w in sr.workloads %}
workload: {{ w.workload }} ({{ w.kind }})
{% if w.recommended %}
{{ w.recommended.variant_key }}
pass {{ w.recommended.pass_rate|pct }} ({{ w.recommended.n_passed }}/{{ w.recommended.n_valid }}) · {% if kind.endswith("cost_usd") %}{{ w.recommended.objective|cost }}{% elif kind == "wall_time_seconds" %}{{ w.recommended.objective|seconds }}{% elif kind == "llm_calls" %}{{ w.recommended.objective|int }} calls{% else %}{{ w.recommended.objective|int }} tokens{% endif %} · tokens {{ w.recommended.median_tokens|int }} · {{ w.recommended.median_wall_time|seconds }}
{% if w.runner_up %}
runner-up: {{ w.runner_up.variant_key }}
{% endif %} {% if w.holdout %}
holdout ({{ w.holdout.task_keys|join(", ") }}): pass {{ w.holdout.pass_rate|pct }} over {{ w.holdout.n_valid }} runs
{% endif %} {% else %}
none eligible
no configuration met the requirement{% if w.best_effort %} · best effort {{ w.best_effort.variant_key }} at {{ w.best_effort.pass_rate|pct }}{% endif %}
{% endif %}
{% endfor %}
{% for w in sr.workloads %}{% if w.vs_runner_up or w.vs_baseline %}

Evidence for {{ w.workload }}: recommended configuration vs alternatives

{% if w.vs_baseline %}{% set p = w.vs_baseline %}{% include "partials/paired.html" %}{% endif %} {% if w.vs_runner_up %}{% set p = w.vs_runner_up %}{% include "partials/paired.html" %}{% endif %} {% endif %}{% endfor %} {% for w in sr.workloads %}
{{ w.workload }} {{ w.configs|length }} configurations · ✔ eligible · ◆ Pareto front (pass rate vs {{ kind }}) {% for f in sr.factor_effects|map(attribute="factor")|unique %}{% endfor %} {% for c in w.configs %} {% for f in sr.factor_effects|map(attribute="factor")|unique %}{% endfor %} {% endfor %}
configuration{{ f }}passvalid{{ kind }}tokenswalltoolsfilesnote
{% if c.eligible %}✔{% elif c.variant_key in w.pareto %}◆{% endif %} {{ c.variant_key }}{{ c.factors.get(f, "") }}{{ c.pass_rate|pct }} {{ c.n_valid }}/{{ c.n_total }} {% if kind.endswith("cost_usd") %}{{ c.objective|cost }}{% elif kind == "wall_time_seconds" %}{{ c.objective|seconds }}{% else %}{{ c.objective|int }}{% endif %} {{ c.median_tokens|int }} {{ c.median_wall_time|seconds }} {{ c.median_tool_calls|int }} {{ c.median_files_changed|int }} {{ c.reason or "" }}
{% endfor %} {% if sr.factor_effects %}

Factor effects (marginal means over configurations)

{% for e in sr.factor_effects %} {% endfor %}
factorlevelconfigsmean pass ratemedian {{ kind }}
{{ e.factor }}{{ e.level }}{{ e.n_configs }}{{ e.mean_pass_rate|pct }}{% if kind.endswith("cost_usd") %}{{ e.median_objective|cost }}{% elif kind == "wall_time_seconds" %}{{ e.median_objective|seconds }}{% else %}{{ e.median_objective|int }}{% endif %}
{% endif %}