Metadata-Version: 2.5
Name: oak-forecast
Version: 0.2.0
Summary: The shared forecasting instrument for the OakQuant platform: overlap-aware evaluation, baseline boards, purged walk-forward, a trials ledger, deflation, and hierarchical partial pooling — domain-free, and held that way by a purity gauge.
Project-URL: Homepage, https://github.com/oakquant-ai/oak-forecast
Author-email: Pumulo Sikaneta <pumulo@gmail.com>
License-Expression: Apache-2.0
License-File: LICENSE
Requires-Python: >=3.13
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == 'dev'
Description-Content-Type: text/markdown

# oak-forecast

**The shared forecasting instrument for the OakQuant platform.** Overlap-aware
evaluation, baseline boards, purged walk-forward, a trials ledger, and deflation
for multiple testing — with no idea what it is scoring.

```bash
pip install oak-forecast
```

## Why it exists

Every domain on this platform that makes a dated, settleable prediction has to be
judged the same way. That machinery was built inside the investments domain,
because that is where the first question was asked — but it was never about
equities. Measured by AST on 2026-09-16, with comments and docstrings stripped so
the scan read code rather than its own explanation:

```
adapt/evaluate    11 files    equity identifiers: 0    equity imports: 0
```

Zero. Every mention of "stock" in it was prose telling the origin story. So the
extraction was a **move, not a rewrite**, and the acceptance test said so out
loud: `adapt evaluate --gate0` had to produce a byte-identical report before and
after — identical *modulo* `generated_at`, which a control run proved is the only
leaf that differs between two runs of unchanged code.

## The parts

| module | what it answers |
|---|---|
| `corpus` | the settled bets, with the label window **measured**, not assumed |
| `overlap` | effective sample size — why `n` is not the denominator |
| `bootstrap` | date-block resampling; the interval that gets quoted |
| `purge` | purged, embargoed walk-forward; the leak in a naive harness |
| `baselines` | what an improvement has to beat before it is one |
| `trials` | how many times we have already looked |
| `deflate` | pricing that search into the significance |
| `report` | the assembled answer, and the gate |
| `host` | the two things the **host programme** supplies |

## The seam

⚠️ **Nothing in this package may know what it is scoring**, and that is enforced
by a purity gauge over the whole package rather than by convention.

Two things legitimately *do* depend on the host: where its frozen corpus lives,
and what its candidate models are. Those are registrations, not imports:

```python
from oak_forecast.evaluate import register_fixture_provider, register_predictor_provider

register_fixture_provider(lambda: "tests/fixtures/my_corpus.json.gz")
register_predictor_provider(lambda: {"my_model": build_my_model})
```

A missing registration **raises**. It does not quietly score nothing — a purged
re-score with no candidates in it exits 0 and reads exactly like "nothing to
worry about", which is the shape of every silent loss this platform has shipped.

## Honesty commitments

1. Every number carries its denominator and its interval.
2. Pre-register before searching — `record(name=..., hypothesis=...)` refuses a
   trial with no hypothesis.
3. A null result is a deliverable.
4. A gate that cannot say "not applicable" trains everyone to ignore it; a gate
   that says "not applicable" must not be read as a pass.

Apache-2.0.
