// backtest validation, not a strategy

Stop trusting a
single backtest run.

A small, dependency-light Python library that runs three checks each caught a real, shipped bug in a live quant-research project — before you find out the hard way that yours didn't.

30 tests 5 / 5 self-mutations caught 0 deps beyond numpy + pandas MIT licensed
$ pip install anchortest
read the incidents →

incident log

Three bugs that looked completely honest

The code ran. The number was real. The chart was plausible. Each one shipped as a "confirmed result" before something caught it.

Incident No. 01 — Anchor.py ✓ caught by anchor_average()

The single-anchor illusion

A strategy rebalanced every 30 trading days; its backtest picked dates the obvious way — every 30th day, starting from wherever the data happened to begin. That "wherever it begins" is a phase, not a neutral choice. Sweeping all 30 possible phases of the cycle found Sharpe ranging from 0.35 to 1.46 — and the one phase every earlier test had used was the single best of the 30.

phases swept: 30 / 30 Sharpe range: 0.35 – 1.46 "lucky" anchor used before: best of all 30
Incident No. 02 — Compare.py ✓ caught by compare()

The sign bug that only breaks one way

Drawdown is stored negative — a shallower 20% drawdown is -0.20. Writing the obvious comparison, candidate_dd <= baseline_dd, silently treats a DEEPER, worse drawdown as a pass: -0.30 <= -0.20 is True. Three candidates shipped as "confirmed improvements" on this exact line before a manual re-check caught it.

false "confirmed" candidates: 3 chars in the bug: 2 (<=) fix: direction handled once, in compare()
Incident No. 03 — Artifacts.py ✓ caught by check_blend_artifact()

The result that had no strategy in it

Blending several phase-shifted "tranches" of one strategy looked like a genuine diversification win: Sharpe up 27%, drawdown several points shallower. The control that caught it: the exact same blend, run on plain buy-and-hold — which cannot have any real phase-dependent skill — "improved" by almost the same ratio. It was smoothing variance across offset windows, not reducing risk.

strategy Sharpe: +27% (looked real) buy-and-hold control: 0.89 → 1.20 (same trick, no skill) would've shipped as: best result of the project

what's in the box

Four functions, each named after what it stops

~400 lines total, on purpose. The value isn't the amount of code — it's that these specific mistakes don't get to cost you what they cost this project.

anchor_average(backtest_fn, params, cycle_days)

Runs your backtest at every phase of its own rebalance cycle and aggregates mean / median / min / worst-case Sharpe — instead of the one phase you happened to write first.

compare(baseline, candidate) / clears_bar(delta)

Per-metric deltas with drawdown's sign convention fixed once. The strict bar requires every tracked metric to improve — "5 of 6" has, in practice, been about a coin flip.

check_blend_artifact(blend_fn)

Runs your tranching / ensembling logic against a phase-invariant synthetic control first. If it "improves" a series with no real skill in it, you've found an artifact.

run_mutation_suite(mutations, test_command)

Deliberately breaks your code in named ways and confirms your own test suite would actually catch each one — a green suite is a claim, not a fact, until you've tried this.


quickstart

Ten lines to a straight answer

# pip install anchortest   (or: pip install -e . from the repo)

from anchortest import anchor_average, compare, clears_bar

def backtest(params, start_shift=0):
    # your backtest. MUST drop start_shift rows of price data before computing anything.
    ...  # returns a pandas Series equity curve, starting at 1.0

baseline, _ = anchor_average(backtest, baseline_params, cycle_days=25, n_anchors=25)
candidate, _ = anchor_average(backtest, candidate_params, cycle_days=25, n_anchors=25)

delta = compare(baseline, candidate)
if clears_bar(delta):
    print("every tracked metric improved -- worth a closer look")
else:
    print("not a promotion:", delta)

work with me

Get your own strategy audited

Before AnchorTest was a library, it was three case files from one real project. I'll run the same checks against YOUR backtest and send back a written report: what passed, what didn't, and exactly why — including when nothing passes.

Quick Check — $199

Anchor-average your existing backtest, run it against the strict multi-metric bar, check any blending/ensembling logic for the artifact pattern above. Written report, 3 business days.

Full Audit — $749

Everything in Quick Check, plus a held-out / out-of-sample test designed specifically for your strategy and universe, plus a 30-minute call to walk through the findings. 7 business days.

book an audit →
grant02339@gmail.com

This is a review of your validation PROCESS — not investment advice, not a recommendation to trade anything, and not a claim that any strategy is profitable. You get an honest report on what the checks show, nothing more.


philosophy

What this is actually for