§2–§3 — GPT-5.6 role benchmarks and seat decisions: landed record (2026-07-17)
Six seats measured for ≈$36 (≤$40 ceiling); one budget-forced reduction (compose n=20, logged).
Verdicts locked by the lead as one batch (seat-switches decision) and applied in
imas-codex ce0144e1. Full measured tables: imas-codex docs/sn-role-benchmarks.md;
raw reports under ~/.local/share/imas-codex/benchmarks/sn_rolebench_*.json.
Verdicts
| Seat | Incumbent | Verdict | Decisive evidence |
|---|---|---|---|
sn-refine | gpt-5.5 | HOLD | Defect resolution 0.874 vs Terra 0.837 / Luna 0.765; lowest collateral rewording (0.147); Luna worse on every quality axis at half cost. |
| names breaker | gpt-5.5 | SWITCH → gpt-5.6-luna | More independent of the blind pair (ρ 0.252 vs 0.319; vs qwen 0.026), better verdict-flip quality (0.714 vs 0.545), ~¼ cost. Flip-quality on 7 flips is the soft axis; independence and cost are the robust n=30 signals. |
| docs breaker | gpt-5.5 | HOLD | Incumbent both most independent (ρ 0.176) and best on overrides (0.286 vs 0.071/0.158); 5.6 tiers flip 2–3× more, nearly always wrongly. |
sn-docs generation | sonnet-4.6 | SWITCH → gpt-5.6-luna | Zero banned-prose findings vs incumbent 15% under the strict-normative prompt (decisive for the campaign's convergence gate); rubric parity 0.813 vs 0.806; 23% cheaper. Sol +0.03 rubric at 2.6× Luna's cost — rejected. |
| domain classifier tier-1 | gpt-5.4 | HOLD | Head-to-head on the 209-path gold set: gpt-5.4 82.3% / Luna 79.4% / haiku-4.5 75.1%, all ≤$0.0006/item. Incumbent wins outright; Luna's two runs (77.5%, 79.4%) also calibrate ~2 pts of run variance, supporting the compose tie-band. |
sn-compose | deepseek-v4-flash (local) | HOLD | n=17: gpt-5.5 0.891 / Terra 0.874 / Luna 0.826 / DSv4 0.814 — Luna↔DSv4 gap 1.2 pts = tie at this sample; DSv4 is $0.00/name marginal (2× H200 serve) and its compose-time descriptions judge above gpt-5.5 (0.957 vs 0.932). kimi-k3 unavailable (launch-day upstream 429s) — re-probe later. |
Bench-harness fixes shipped en route
0a05c173— refine bench starved the prompt of the closed-vocabulary grammar listing production injects (candidates invented tokens).0baa04ee— refine bench passed no scored examples, the prompt's onlykindsignal; plus a loud-fail tripwire (n=0 raises with a failure breakdown instead of saving a hollow report).6ef0ab2c— the latent fragility fixed at source: kind vocabulary unified on the ISN catalogKindenum (LinkML trimmed, LLM schemas carry a flat in-schema enum, prompts enumerate kinds; three-way consistency test). Format correctness no longer depends on examples.7c846191— local endpoints for models appearing in review-quorummodelslists (unblockedsn bench --models hosted_vllm/deepseek-v4-flash).cd18c024— pipeline-version clear gate retired (drift logs, never blocks; the wipe recommendation is gone).
Campaign cost input (feeds §5)
Docs generation at Luna's measured ≈$0.09/doc with a 0% generator banned-prose rate; refine stays gpt-5.5 at ≈$0.08/refine. The §5 pilot supplies the end-to-end per-name number.