DashboardProbes › Feature-Realness Instrument-Grounding Loop

opendeviationbar-py · bounded campaign · started 2026-07-22

🧪 Feature-Realness Instrument-Grounding Loop

Before a promoted feature can be called real or useful, the instrument that judges it must itself be proven on data where the right answer is known. This bounded loop grounds the 20 candidate instruments (#0–#19) — one per iteration, in dependency order — then STOPS. Every iteration emits one HTML page, appended to the ledger at the bottom of this index.

3 / 20
instruments grounded
17
discovered (candidates)
0 / 30
columns with grounded evidence
P2 ▸
#2 grounded · #3 next

What this loop does

Each candidate instrument sits an entrance exam: the §4 five-control battery (positive / negative-null / substitution / power-calibration / invariance-leakage), plus its instrument-specific Harden item and a published detection envelope. It ends with a terminal verdict:

DISCOVEREDUNDER EVALUATIONGROUNDED (or REJECTED / UNVERIFIABLE-PARK)

Only a GROUNDED instrument may later judge a real feature. The graduation tally is the loop's STOP condition: when all 20 carry a terminal verdict, the loop publishes a completion page and ends.

Substrate gate. Instruments #1 (future-perturbation invariance) and #0 (CPCV purge/embargo) ground first. If either fails to ground, the campaign HALTS before the realness/utility instruments — every downstream out-of-sample number rides on the substrate.

Hard operating rules

Dependency order (frozen)

PhaseInstruments (in order)
SEALreal-data-only control substrate self-validation (first firing)
Substrate#1 future-perturbation invariance → #0 CPCV purge/embargo (both GROUNDED or HALT)
Realness#2 block-perm null → #3 effective-n+FDR → #6 SFI → #8 DSR → #9 PBO → #10 Harvey-Liu-Zhu
Usefulness#4 Rank IC → #5 IC-decay → #7 MDA → #19 quantile monotonicity
Conditional#11 CMI → #12 knockoffs → #13 DML
Increment / mechanism#14 spanning → #15 mechanism-intensity
Robustness / detector#16 structural-break → #17 plateau → #18 detector lead-lag

Full per-instrument Admit / Graduate / Harden rows: METRIC-EVALUATION-FRAMEWORK.md §7. Metric tables twin: The Realness Question.

Next iteration

#3 — effective-n deflation + BH-FDR (realness / multiple-testing). #2 grounded (iter 5) — the first realness instrument is cleared. Next: ground #3 — the HAC long-run-variance effective-n estimator + Benjamini-Yekutieli FDR (autocorrelated data has fewer independent points than it looks; count them honestly, then correct for many tries), with the mandatory GARCH heavy-tail null (the archetype the scalar effective-n deflator fails) and iid/AR1/white/block null calibration. Then #6 → #8 → #9 → #10.

Iteration ledger (append-only)

One row per firing. The loop appends below the marker; rows are never edited or deleted (corrections are new rows with supersedes:). Each row links the iteration's own HTML page.

#DateInstrument / stepVerdictPage
12026-07-22SEAL · real-data-only control-substrate self-validationSEAL-PASS lab openiter 1 →
22026-07-22#1 · future-perturbation invariance (causality floor)GROUNDEDiter 2 →
32026-07-22#0 · CPCV purge/embargo (leakage-free partition)CHECKPOINT 6/7 · PBO openiter 3 →
42026-07-22#0 · CPCV purge/embargo (supersedes iter 3) — ★ substrate gate PASSGROUNDEDiter 4 →
52026-07-22#2 · block-permutation shuffled-label null (realness) · P2GROUNDEDiter 5 →