{# Paired statistical evidence for "B vs A". Expects `p` (PairedComparison). #} {% if p %}

{{ p.b }}: {{ p.verdict }} vs {{ p.a }} over {{ p.n_tasks }} paired task{{ '' if p.n_tasks == 1 else 's' }} (wins {{ p.wins }}, losses {{ p.losses }}, ties {{ p.ties }}{% if p.sign_test_p is not none %}, sign test p = {{ '%.3f'|format(p.sign_test_p) }}{% endif %})

{% for label, iv, unit in [("pass rate", p.pass_rate_diff, "pts"), (p.cost_kind, p.cost_diff, ""), ("llm_calls", p.llm_calls_diff, ""), ("improvement ratio", p.improve_ratio_diff, "")] if iv is not none or label != "improvement ratio" %} {% if iv %} {% if unit == "pts" %} {% else %} {% endif %} {% else %} {% endif %} {% endfor %}
per-task difference (B − A)mean{{ ((p.pass_rate_diff.level if p.pass_rate_diff else 0.95) * 100)|round|int }}% intervalP(> 0)
{{ label }}{{ (iv.estimate * 100)|delta(1) }} pts [{{ (iv.low * 100)|delta(1) }}, {{ (iv.high * 100)|delta(1) }}]{{ iv.estimate|delta(4) }} [{{ iv.low|delta(4) }}, {{ iv.high|delta(4) }}]{{ iv.p_positive|pct }}———
{% if not p.enough_tasks %}

Fewer than {{ p.min_tasks }} paired tasks: the interval is shown, but no verdict is given. Add tasks (repetitions do not count as tasks).

{% endif %}
{% endif %}