stabreg_kappa_sweep_m3.py)Key insight that made it cheap: e_total and oos_mse per feature-set are ฮบ-independent โ the knob only enters at survivor selection. So one 130-subject ร 8-set grid computation (identical machinery to the entrance exam, per-fold MSEs retained) supports the entire ฮบ-sweep and the M3 census post-hoc.
| Fight | Reading | Ruling |
|---|---|---|
| ฮบ knob (ยง0: derive or show insensitive) | Derived ฮบ_D = 0.010364 โ the median (over 130 subjects) of the estimation-noise SE of the best set's fold MSEs. The 0.01 exam fixture sat almost exactly on the derived value. Plateau: adjacent-ฮบ readout Spearman 0.948โ1.0 across the ENTIRE sweep [0 โ 0.5]; local insensitivity at ฮบ_D/2..2ฮบ_D: median ฮn_surv = 0. Not degenerate: trivial-family rate โค 0.846 < 0.95 (at ฮบ_D: 69 empty ยท 38 all-8 ยท 23 non-trivial). | RESOLVED โ dial derived AND shown insensitive; the iter-37 red ink dissolves |
| M3 redundancy attack (kill bar |ฯ| โฅ 0.95) | Readout 1, n_surv: vs e-ICP-fp n_acc ฯ = 0.977 โ (vs eicp_log_min_e โ0.877 ยท ipp_n_acc 0.641 ยท seqicp 0.595 ยท ias 0.118 ยท k_v 0.211 ยท mean-redundancy โ0.315). Readout 2, w_sum: worst 0.804 โ same instrument. Rule kills on worst across readouts. | FAIL โ 0.977 โฅ 0.95 |
| Verdict (row 51) | EXCLUDE(redundant โ set-level readout repackages e-ICP-fp accepted-set count, ฯ=0.977) โ BONEYARD. Re-proposal admissible only with a stability screen NOT built on the admitted v5 oracle. | |
WHY THE KILL WAS PREDICTED (iter 37, "shared-oracle question"):
StabReg survivor = passes STABILITY screen AND passes PREDICTIVENESS screen
โโโ the admitted v5 โโโ adds too little independent
fold-product e-value variation at census scale
oracle ITSELF
โ
โผ
count(surviving sets) โ count(e-value-accepted sets) = e-ICP-fp's n_acc โ ฯ = 0.977