Forecasts chosen by what would have worked.
A forecasting package that replays the past before trusting a model: rolling-origin backtest, choice by out-of-sample error, and intervals taken from the errors actually observed. The Rust crate foresight inside: fast, on all cores, with no runtime dependencies.
Built to be believed, not just to fit
A model that fits the past well can still forecast badly. Every choice here is made on forecasts the model gave without seeing the answer.
Rolling-origin backtest
Each model is refitted at every origin using only earlier data and forecasts 1 to h periods ahead. Errors are reported by horizon: MAPE, MAE, RMSE, MASE and bias.
Choice by out-of-sample error
The winner is whatever forecast best, the simple average of the best models included. No information criterion decides what you ship.
Empirical intervals
Intervals are quantiles of the errors seen in the backtest, horizon by horizon. No normality is assumed and none is needed.
Intervals for totals
The total of the next k periods gets its own interval, measured on totals. Adding up monthly limits would overstate the uncertainty of a year.
Rust inside
The models and the backtest are the Rust crate, compiled into the package: the backtest runs on all cores and releases the GIL. Lists, NumPy arrays and pandas Series go in; pandas comes out on request.
Deterministic
The same series gives the same numbers, on one core or sixteen. Fitted parameters are exposed by name, ready for an audit trail.
Models
From the yardsticks every model must beat to the exponential smoothing family, seasonal ARIMA with automatic orders, Prophet with dated events, TBATS for several seasonal periods and Croston for intermittent demand. Models can be combined in an ensemble or run on a series decomposed by STL. Any of them can run on the log or another Box-Cox scale.
The Python package of the Rust crate foresight. ARIMA is estimated by exact Gaussian maximum likelihood with the innovations algorithm; Prophet is fitted without Stan, its changepoints that do not matter coming out as exactly zero. Any model can be renamed with named() and combined with others in an Ensemble.
On real data
Monthly ICMS, the sales tax of the Brazilian state of Piauí, from public fiscal reports. The average of two seasonal ARIMA models, on the log and on the original scale, had a mean absolute percentage error of 3.4% over 36 origins and 12 horizons; the seasonal naive forecast had 11.5%.
ICMS revenue, Piauí: last four years and the next twelve months
BRL million per month. Forecast from data up to June 2026, with the 80% empirical interval.
Forecast as a table
| Month | Forecast | Lower (80%) | Upper (80%) |
|---|
Total of the next six months: BRL 5.02 billion, 80% interval 4.88 to 5.25 billion. Adding up the six monthly limits instead would give a wider and wrong interval.
Source: Siconfi/STN, RREO Anexo 03. The data is in tests/data/ of the repository.
How it works
One pass over the past produces everything: the ranking, the choice and the intervals.
Describe the series
Values and seasonal period. Slices keep season and position.
Pick candidates
The built-in set, or your own list with your own names.
Replay the past
Origins, horizon and window are yours to set.
Read the report
Errors by horizon, the choice, forecasts and intervals.
Install & run
Install, describe the series, run the backtest. Python 3.9 or later; wheels for Linux, macOS and Windows, imported as foresight.
$ pip install pyforesight
# choose among every built-in model
import foresight as fs
y = fs.monthly(values)
report = fs.backtest(y)
best = report.best
for p in best.forecast:
lo, hi = p.interval(0.80)
print(p.horizon, p.mean, lo, hi)
total = best.cumulative(6)
report.to_pandas()
# one model on its own
fit = fs.Theta().fit(values)
nxt = fit.forecast(3)
# your own settings and your own names
report = fs.backtest(y, [
fs.SeasonalNaive().named("same_month"),
fs.log(fs.Arima.airline()),
fs.Ensemble(fs.defaults(), weighting="stacked"),
], origins=24, horizon=6, window=60)
Architecture
One implementation in Rust. The Python package, the R package and Rust programs call it, so the numbers are the same everywhere; the Go edition is an independent port, checked against it.
The numbers of the Rust crate
The package runs the Rust crate, so its numbers are the crate's. The tests check that nothing is lost on the way; the crate in turn is compared with R and reproduced independently in Go.
| Against | What | Agreement |
|---|---|---|
| Results recorded by the Rust crate | ARIMA, regression with ARIMA errors, ETS, Prophet with events, TBATS, STL and MSTL, Croston, cleaning, ensembles, tests of stationarity and seasonality, backtests of 11 and 18 candidates on three public series | The same numbers, up to the floating point of each platform |
R, forecast 9.0.2 and prophet 1.1.7 | The crate: benchmarks, Theta, ARIMA, ETS, STL, TBATS, Croston, Prophet | Benchmarks, STL and Croston exact; ARIMA within 0.01%; the likelihood of ETS and TBATS never worse than R's; Theta within 0.1% and Prophet within 0.5% |
| foresight-go | An independent implementation in Go of the same methods | The same choices; the same numbers within the precision of each search |
CI on every platform
rustfmt · clippy · tests on Linux, macOS and Windows, with Python 3.9 and the current release; type stubs checked against the module.
From the papers
Box & Jenkins, Brockwell & Davis, Hyndman & Khandakar, Taylor & Letham, Assimakopoulos & Nikolopoulos, Guerrero. R supplies reference numbers only.
MIT, nothing attached
No runtime dependencies: nothing to reconcile, no supply chain to watch beyond the crate itself.