§2–§3 — GPT-5.6 role benchmarks and seat decisions: landed record (2026-07-17)

Six seats measured for ≈$36 (≤$40 ceiling); one budget-forced reduction (compose n=20, logged). Verdicts locked by the lead as one batch (seat-switches decision) and applied in imas-codex ce0144e1. Full measured tables: imas-codex docs/sn-role-benchmarks.md; raw reports under ~/.local/share/imas-codex/benchmarks/sn_rolebench_*.json.

Verdicts

SeatIncumbentVerdictDecisive evidence
sn-refinegpt-5.5HOLDDefect resolution 0.874 vs Terra 0.837 / Luna 0.765; lowest collateral rewording (0.147); Luna worse on every quality axis at half cost.
names breakergpt-5.5SWITCH → gpt-5.6-lunaMore independent of the blind pair (ρ 0.252 vs 0.319; vs qwen 0.026), better verdict-flip quality (0.714 vs 0.545), ~¼ cost. Flip-quality on 7 flips is the soft axis; independence and cost are the robust n=30 signals.
docs breakergpt-5.5HOLDIncumbent both most independent (ρ 0.176) and best on overrides (0.286 vs 0.071/0.158); 5.6 tiers flip 2–3× more, nearly always wrongly.
sn-docs generationsonnet-4.6SWITCH → gpt-5.6-lunaZero banned-prose findings vs incumbent 15% under the strict-normative prompt (decisive for the campaign's convergence gate); rubric parity 0.813 vs 0.806; 23% cheaper. Sol +0.03 rubric at 2.6× Luna's cost — rejected.
domain classifier tier-1gpt-5.4HOLDHead-to-head on the 209-path gold set: gpt-5.4 82.3% / Luna 79.4% / haiku-4.5 75.1%, all ≤$0.0006/item. Incumbent wins outright; Luna's two runs (77.5%, 79.4%) also calibrate ~2 pts of run variance, supporting the compose tie-band.
sn-composedeepseek-v4-flash (local)HOLDn=17: gpt-5.5 0.891 / Terra 0.874 / Luna 0.826 / DSv4 0.814 — Luna↔DSv4 gap 1.2 pts = tie at this sample; DSv4 is $0.00/name marginal (2× H200 serve) and its compose-time descriptions judge above gpt-5.5 (0.957 vs 0.932). kimi-k3 unavailable (launch-day upstream 429s) — re-probe later.

Bench-harness fixes shipped en route

Campaign cost input (feeds §5)

Docs generation at Luna's measured ≈$0.09/doc with a 0% generator banned-prose rate; refine stays gpt-5.5 at ≈$0.08/refine. The §5 pilot supplies the end-to-end per-name number.