{% set ar = ablation_report %}

Component ablation bundle {{ ar.bundle_hash[:12] }} on {{ ar.base_variant }} · every component tested against its own absence

{% for note in ar.notes %}

{{ note }}

{% endfor %}

Whole bundle: full vs minimal (no components)

{% set p = ar.full_vs_minimal %} {% include "partials/paired.html" %} {% for e in ar.components %} {% set c = e.comparison %} {% endfor %}
componentverdictΔ pass rateintervalΔ {{ ar.full_vs_minimal.cost_kind }}Δ llm_callswins / losses / tiessign test p
{{ e.component }} {{ e.verdict }} {% if c.pass_rate_diff %}{{ (c.pass_rate_diff.estimate * 100)|delta(1) }} pts{% else %}—{% endif %} {% if c.pass_rate_diff %}[{{ (c.pass_rate_diff.low * 100)|delta(1) }}, {{ (c.pass_rate_diff.high * 100)|delta(1) }}]{% else %}—{% endif %} {% if c.cost_diff %}{{ c.cost_diff.estimate|delta(4) }}{% else %}—{% endif %} {% if c.llm_calls_diff %}{{ c.llm_calls_diff.estimate|delta(2) }}{% else %}—{% endif %} {{ c.wins }} / {{ c.losses }} / {{ c.ties }} {% if c.sign_test_p is not none %}{{ '%.3f'|format(c.sign_test_p) }}{% else %}—{% endif %}

Effect = full bundle minus the bundle without that component, per task. A component "helps" only when the interval excludes zero over enough tasks; "no evidence" is a reason to prefer the simpler harness.