iteration 23 · 2026-07-23 · conditional axis · bounded slice 2 · laptop-drives-bigblack
🎛️ #12 Model-X Knockoffs — slice 2 CHECKPOINT
A verify-before-report course-correction. Slice 1 blamed the null-FDR inflation on row-autocorrelation (the Harden's "#1 gaming"). Slice 2 systematically falsified that — AR whitening, fractional differencing, spacing and block-size all leave FDR ~0.2–0.3 — and traced the real cause to the selection rule: plain-knockoff is anti-conservative; knockoff+ controls FDR (0.0) regardless of autocorrelation, needing only a wide-enough panel. The interaction-only non-linear statistic is cleanly grounded.
0.0
knockoff+ null-FDR (was 0.25)
RF 1.0
interaction W (lasso 0.0)
0.64
knockoff+ power (wide panel)
slice 2
wide-panel + k+ next
Preflight (resource-only): load1 1.67 · 38 GiB available · si/so ~0 · ClickHouse active, readonly=2. 5c/5G/no-swap capped; single-thread BLAS.
Grounded machinery (SOLID)
| Gate | Result (seed 20260723) | |
| Exchangeability 2nd-moment | 0.028 (near-singular ρ=1 pair floor) | PASS |
| Power (dense main-effect) | 0.98, derand Π 1.0 | PASS |
| Interaction-only non-linear W | RF-W 1.0 selects the sign-XOR feature every rep; lasso-W 0.0 misses every rep | PASS |
| Attribute @ρ=.80 | signal-carrier A 1.0, decoy B 0.13 | PASS |
The interaction result is the whole point of #12 on the conditional axis: a non-linear (RandomForest) importance-difference statistic selects the interaction-only feature that rank-IC (#4), IC-decay (#5), MDA (#7) and monotonicity (#19) all flag FNR=1 — while the linear lasso statistic (0.0) confirms the non-linearity is load-bearing, not the knockoff framework alone.
The course-correction — FDR inflation is NOT autocorrelation
Slice 1's hypothesis (row-autocorrelation breaks iid-exchangeability → FDR inflates) was tested against every plausible fix and falsified:
| Construction (null-FDR, target ≤.10) | FDR |
| iid consecutive · block-perm 250 | 0.30 |
| iid consecutive · block-perm 50 | 0.20 |
| spaced ×50 · block 250 / 50 | 0.25 / 0.30 |
| AR pre-whitening | 0.20 |
| uniform fractional differencing (d=0.4) | 0.20 (+ induces spurious AC 0.13→0.48) |
None move the inflation → autocorrelation is not the cause. Fractional differencing was actively harmful: the panel has heterogeneous memory (most features near-white, AC 0.13–0.28), and a uniform wide-window difference over-differences them and induces autocorrelation.
The real cause — the selection rule. Plain-knockoff (Barber–Candès, no +1) controls only a modified-FDR and is anti-conservative. knockoff+ (with the +1 numerator) gives null-FDR 0.0 and controls FDR regardless of the autocorrelation — but the +1 makes it unable to reject fewer than 1/q features (the detection floor), so it needs a wide panel. On p=20 (8 real + 12 planted): knockoff+ controls FDR with power 0.64 @IC=.03, vs plain's 0.93 (uncontrolled). The autocorrelation never broke FDR — the wrong selection rule did.
Slice 3 (terminal plan): stage a wide real-feature panel (~25 features) + MVR + Ledoit–Wolf + knockoff+ / derandomised e-BH (FDR-controlled and more powerful than plain knockoff+); complete power/FDP & null-FDR-with-power, and re-confirm interaction-only + attribute/ABSTAIN under the FDR-controlling rule. The Harden's ∀H≤.85 becomes "verify knockoff+ FDR control across the natural feature Hurst" — empirically it holds — not a whitening requirement.