what's in the box
~400 lines total, on purpose. The value isn't the amount of code — it's that these specific mistakes don't get to cost you what they cost this project.
Runs your backtest at every phase of its own rebalance cycle and aggregates mean / median / min / worst-case Sharpe — instead of the one phase you happened to write first.
Per-metric deltas with drawdown's sign convention fixed once. The strict bar requires every tracked metric to improve — "5 of 6" has, in practice, been about a coin flip.
Runs your tranching / ensembling logic against a phase-invariant synthetic control first. If it "improves" a series with no real skill in it, you've found an artifact.
Deliberately breaks your code in named ways and confirms your own test suite would actually catch each one — a green suite is a claim, not a fact, until you've tried this.
quickstart
# pip install anchortest (or: pip install -e . from the repo) from anchortest import anchor_average, compare, clears_bar def backtest(params, start_shift=0): # your backtest. MUST drop start_shift rows of price data before computing anything. ... # returns a pandas Series equity curve, starting at 1.0 baseline, _ = anchor_average(backtest, baseline_params, cycle_days=25, n_anchors=25) candidate, _ = anchor_average(backtest, candidate_params, cycle_days=25, n_anchors=25) delta = compare(baseline, candidate) if clears_bar(delta): print("every tracked metric improved -- worth a closer look") else: print("not a promotion:", delta)
work with me
Before AnchorTest was a library, it was three case files from one real project. I'll run the same checks against YOUR backtest and send back a written report: what passed, what didn't, and exactly why — including when nothing passes.
Anchor-average your existing backtest, run it against the strict multi-metric bar, check any blending/ensembling logic for the artifact pattern above. Written report, 3 business days.
Everything in Quick Check, plus a held-out / out-of-sample test designed specifically for your strategy and universe, plus a 30-minute call to walk through the findings. 7 business days.
grant02339@gmail.com
This is a review of your validation PROCESS — not investment advice, not a recommendation to trade anything, and not a claim that any strategy is profitable. You get an honest report on what the checks show, nothing more.
philosophy
Compare against a control before you trust a transform. If you can build a version of your data where the true answer is known, run your method on that first. Blend-artifact checking is one instance of this; it generalizes further than this library currently automates.
Partial improvement has been about a coin flip. A candidate that improves mean Sharpe while its worst-case or recent-half Sharpe gets worse is not free money — that exact pattern has predicted a real regression as often as a win.
A green test suite is a claim, not a fact.
run_mutation_suite is the smallest version of "did this test assert anything, or did
it just not crash" — writing this library's own tests, it caught two bugs before they shipped.