Forward-Orthogonality · feature-declaration expiry · robustness review
Does a feature declared orthogonal / conditional today have a trustworthy way to know when that verdict expires in live trading? This reviews the design we already have (the G12 live monitor + the feature-status lifecycle) against what we need now — robustness as research.
2026-07-15 Existing design: EVALUATING — not operational Benchmark: campaign laws + SOTA
A feature is graded per regime through an early-exit gate ladder; which gate first disqualifies it sets the grade. Promoted features (CONDITIONAL/STABLE) ship with a regime scope and a live expiry monitor. When the monitor fires, the verdict expires and the feature re-runs the sealed ladder.
feature × regime ──► GATE LADDER (G0 orthogonal-today · G3 regime-invariance ·
G2 forward-lift · G10 shuffled-null · G11 e-BH)
│ first gate to disqualify = the grade
┌───────────┬──────────┴──────────┬────────────────────┐
▼ ▼ ▼ ▼
FAIL FRAGILE CONDITIONAL STABLE-ORTHOGONAL
(=BAN) (=WATCH) (=WATCH) (=ORTHOGONAL)
└────────── promoted, scoped to regimes R ──────┘
│
▼
G12 / E6 LIVE MONITOR (post-promotion, streaming)
CTM/SKIT wealth martingale + BOCPD alarm, e-BH-merged,
Ville bound: P(∃t: W_t ≥ 1/α) ≤ α → "verdict EXPIRED"
│ fires (or: next completed regime slice)
▼
EXPIRE → feature reverts to DRAFT → re-run sealed ladder
→ re-promote (new scope) OR demote
Grounded grading evidence (strong): persistence 94–95% into next regime · flip-ranking AUC 0.95/0.88 (crypto/forex) · ACI coverage 89.9%/79.9% · FDR-declared sets at q=0.10, PBO 0.00/0.20, deflation passed · ICP 0/92 ⇒ every declaration is regime-conditional · leakage-gated, twice-replicated. Honest ceiling: ~8/10 structural — unconditional invariance is provably impossible.
Read: we can grade a feature and we picked the right kind of alarm — but the alarm itself, the part that says "this verdict has expired," is the weak link, and the automatic re-grounding depends on a detector we haven't built.
CTM + SKIT + e-BH) that the later matrix-admission campaign put under attack — and it proved that any monitor whose "how big is chance?" null assumes independence cries wolf at 11–57% instead of 5% on market bars. The design (June 2026) predates that discovery (July 2026), so its default null is the one the campaign banned. Good news: the campaign also certified the fix (the SKIT betting e-process + circular de-alignment), so the hole is closable — but nothing gates until it passes the certificate attack.
| # | Gap | Breaks | Severity |
|---|---|---|---|
| F1 | Generic CTM fed raw non-conformity scores from autocorrelated bars → 11–57% false-alarm (killed 8 instruments). | L-IID (dependence-aware null) | DECISIVE |
| F2 | E6 replay never run — no positive/negative control; sensitivity + specificity unmeasured. | M5 blind-gauge | DECISIVE |
| F3 | No attribution — feature-drift vs instrument/substrate-drift indistinguishable. | alarm fatigue / spurious re-cert | HIGH |
| F4 | One α-monitor per (feature × regime) × hundreds → family-wide false-alarm explosion; e-BH not wired at the monitor layer. | L-IID / L-MAGIC (aggregation) | HIGH |
| F5 | Interim "re-ground after next completed regime slice" needs a live slice-completed detector — unbuilt (B-05 gated). | auto-expiry inert | HIGH |
| F6 | WATCH covariate weights estimated on ~2K windows; rolling CTM can self-heal toward the new regime. | estimation-error / drift-to-accept | MED |
| F7 | The dependence-aware null's own inputs (block-length / tail-index) can drift; nothing re-checks them. | live L-IID | MED |
| F8 | ~2K bars/regime — effective sample size can be too thin; no under-power veto. | false negatives | MED |
| F9 | wold_R crypto inversion (0.278) · forex PBO watch. (forex ICP wiring defect — later REPAIRED by the admission loop.) | grading hygiene | LOW |
Priority-ordered; each closes a numbered gap above.
feature verdict; an instrument verdict suppresses expiry.Keep: the anytime-valid guarantee, the e-BH aggregation intent, the honest regime-scoped grammar, and the whole well-grounded grading side — these are SOTA-correct.
Fix before it can be trusted live, in order: (1) a dependence-aware null that passes the certificate attack, (2) the E6 positive/negative-control validation, (3) an attribution race so a regime turn doesn't spuriously expire still-valid features. Everything else is structural hardening on top.
The single sentence: our expiry design is right, but its default statistics were written before we learned that market memory breaks naïve tests — so the expiry monitor must be re-nulled for dependence and power-proven (E6) before any feature's verdict is allowed to expire on its say-so.
Research assessment · no code, no production change, nothing ratified · append-only.
SSoT twin: findings/evolution/audits/2026-05-26-forward-orthogonality-prediction/FEATURE-EXPIRY-ROBUSTNESS-ASSESSMENT-2026-07-15.md.
Sources: PHASE-3-PROBE-DESIGN.md (G0–G12) · STATUS-GROUNDING-AND-THRESHOLDS.md §2 · SOTA-RESEARCH-2026-06-23.md §2 · CAMPAIGN-GROUNDED-SUMMARY.md · probes/forward.html · matrix-admission LEDGER rows 99–117 (the two laws) · robust-expiry SOTA workflow wf_88195813-778.
Citation-verification: pre-2026 backbone established; 2026-dated SOTA pending the standing arXiv re-verification gate before wiring.