DashboardProbesRealness Loop › Iter 18 · #19

iteration 18 · 2026-07-23 · usefulness axis · laptop-drives-bigblack

📊 #19 Quantile monotonicity + PT/RW MR CHECKPOINT · in-progress

Does the mean forward return increase monotonically across feature quantile bins (healthy), or spike in one tail (rare-veto)? The Patton–Timmermann / Romano–Wolf monotonic-relation test. The mechanics are right (5/6 base-seed gates), but three genuinely-hard sub-problems on real 3s-return data need careful work before it grounds.

5 / 6
base-seed gates
1.00
power(PT) @ IC=.05
.042
FPR (per-test permutation)
3 open
robustness items
Preflight (resource-only): load1 3.43 (elevated, other users; no heavy nasimubd job) · 42 GiB available · si/so ~0 · ClickHouse active, readonly=2. 5c/5G/no-swap capped.

State across 5 seeds

Gate (§7 row 19)Base seedSeed range
Admit demoPT p 0.01, MR% 0.778MR% 0.556–0.889fragile
Power @ IC=.051.001.00robust
Both nulls FPR ∈ [.03,.07]0.042 / 0.042bp 0.008–0.375heavy-tie
RW trap (rare-veto ≤ .05)plain PT 1.0, RW 0.10RW 0.000–0.167leaks
Collinear decoy collapses1.000.97–1.00robust
Envelope FNR (non-monotone)veto 0.90, inter 1.00.83–1.00robust

What is solid (real progress)

Three open items (genuinely hard on real 3s returns)

  1. Heavy-tie binning. Heavy-tie features (aggregation_density, lookback_hurst) give degenerate quantile bins → FPR corrupts to 0.375 on one seed (the #4 heavy-tie class). Fix: full-spread-only null pool or sub-ULP jitter.
  2. The monotone exemplar is extremes-dominated. Even IC=0.30 gives a bin profile huge at the tails and ≈0 in the middle ([-0.82, ~0×8, +0.80]) — not monotone-throughout — so MR% swings 0.556–0.889 by seed. Fix: a robustly monotone exemplar (IC~0.5 or a real driver).
  3. The rare-veto trap leaks (RW 0.0–0.167; need ≤.05). Fix: require ≥k significant increments spread across the range (a true RW-FWE step-down).
Resolution (next firing, pre-registered): tie-robust binning → robustly-monotone admit exemplar → sharper rare-veto trap → seed sweep, then a terminal verdict. The usefulness axis is already complete (#4, #5, #7); this is the 4th usefulness instrument.