Metadata-Version: 2.4
Name: forestci
Version: 0.9
Summary: Confidence intervals for scikit-learn forest algorithms
Author: Bryna Hazelton, Kivan Polimis
Author-email: Ariel Rokem <arokem@uw.edu>
Maintainer-email: Ariel Rokem <arokem@uw.edu>
License-Expression: MIT
Project-URL: Documentation, https://contrib.scikit-learn.org/forest-confidence-interval/
Project-URL: Changelog, https://github.com/scikit-learn-contrib/forest-confidence-interval/blob/master/CHANGELOG.md
Project-URL: Source, https://github.com/scikit-learn-contrib/forest-confidence-interval
Project-URL: Issues, https://github.com/scikit-learn-contrib/forest-confidence-interval/issues
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.20
Requires-Dist: scikit-learn<2.0,>=1.0
Requires-Dist: scipy
Provides-Extra: test
Requires-Dist: pytest==9.0.3; extra == "test"
Requires-Dist: tqdm; extra == "test"
Provides-Extra: progress
Requires-Dist: tqdm; extra == "progress"
Provides-Extra: docs
Requires-Dist: matplotlib; extra == "docs"
Requires-Dist: numpydoc; extra == "docs"
Requires-Dist: pandas; extra == "docs"
Requires-Dist: pillow; extra == "docs"
Requires-Dist: sphinx; extra == "docs"
Requires-Dist: sphinx-autoapi; extra == "docs"
Requires-Dist: sphinx-gallery; extra == "docs"
Requires-Dist: sphinx-rtd-theme; extra == "docs"
Provides-Extra: dev
Requires-Dist: flake8; extra == "dev"
Requires-Dist: matplotlib; extra == "dev"
Requires-Dist: numpydoc; extra == "dev"
Requires-Dist: pandas; extra == "dev"
Requires-Dist: pillow; extra == "dev"
Requires-Dist: pytest==9.0.3; extra == "dev"
Requires-Dist: sphinx; extra == "dev"
Requires-Dist: sphinx-autoapi; extra == "dev"
Requires-Dist: sphinx-gallery; extra == "dev"
Requires-Dist: sphinx-rtd-theme; extra == "dev"
Requires-Dist: tqdm; extra == "dev"
Dynamic: license-file

# `forestci`: confidence intervals for Forest algorithms

[![Python package](https://github.com/scikit-learn-contrib/forest-confidence-interval/actions/workflows/pythonpackage.yml/badge.svg?branch=master)](https://github.com/scikit-learn-contrib/forest-confidence-interval/actions/workflows/pythonpackage.yml)
[![Documentation](https://github.com/scikit-learn-contrib/forest-confidence-interval/actions/workflows/docbuild.yml/badge.svg?branch=master)](https://contrib.scikit-learn.org/forest-confidence-interval/)
[![status](http://joss.theoj.org/papers/b40f03cc069b43b341a92bd26b660f35/status.svg)](http://joss.theoj.org/papers/b40f03cc069b43b341a92bd26b660f35)

Forest algorithms are powerful [ensemble methods](http://scikit-learn.org/stable/modules/classes.html#module-sklearn.ensemble) for classification and regression. 
However, predictions from these algorithms do contain some amount of error. 
Prediction variability can illustrate how influential
the training set is for producing the observed random forest predictions.

`forest-confidence-interval` is a Python module that adds a calculation of
variance and computes confidence intervals to the basic functionality
implemented in scikit-learn random forest regression or classification objects.
The core functions calculate an in-bag and error bars for random forest
objects.

Version 0.9 supports scikit-learn 1.0 through 1.9, including the signature
changes introduced in scikit-learn 1.9.

This module is based on R code from Stefan Wager 
([`randomForestCI`](https://github.com/swager/randomForestCI) deprecated in favor of [`grf`](https://github.com/swager/grf))
and is licensed under the MIT open source license (see [LICENSE](LICENSE)).
The present project makes the algorithm compatible with [`scikit-learn`](https://scikit-learn.org/stable/).

To get the proper confidence interval, you need to use a large number of trees (estimators). 
The [calibration routine](https://github.com/scikit-learn-contrib/forest-confidence-interval/pull/114) 
(which can be included or excluded on top of the algorithm) tries to extrapolate
the results for an infinite number of trees, but it is instable and it can cause numerical errors:
if this is the case, the suggestion is to exclude it with `calibrate=False` 
and test increasing the number of trees in the model to reach convergence.

## Installation and Usage

Before installing the module you will need `numpy`, `scipy` and `scikit-learn`.

To install `forest-confidence-interval`, run:

```
python -m pip install forestci
```

To install the development version from GitHub, use:

```shell
python -m pip install git+https://github.com/scikit-learn-contrib/forest-confidence-interval.git
```

Usage:

```python
import forestci as fci
ci = fci.random_forest_error(
  forest=model, # scikit-learn Forest model fitted on X_train
  X_train_shape=X_train.shape,
  X_test=X, # the samples you want to compute the CI
  inbag=None,
  calibrate=True,
  memory_constrained=False,
  memory_limit=None,
  y_output=0 # in case of multioutput model, consider target 0
)
```

## Examples

The examples (gallery below) demonstrate the package functionality with random forest classifiers and regression models.
The regression examples use a bundled copy of the
[UCI Auto MPG dataset](https://doi.org/10.24432/C5859H), while the classifier
example simulates how to add measurements of uncertainty to tasks like
predicting spam emails. Keeping the regression data in the repository makes
documentation builds reproducible without relying on an external data service.

[Examples gallery](http://contrib.scikit-learn.org/forest-confidence-interval/auto_examples/index.html)

## Rebuilding and publishing the documentation

Documentation builds run **only when manually requested**, not on pushes or
pull requests. In GitHub, open **Actions → Documentation build → Run workflow**,
select the desired branch under **Use workflow from**, and click **Run workflow**.
For a pull request, select its source branch in this repository, not `master`;
the build checks out and publishes that selected branch. Because every branch
publishes to the same website, a new run cancels any previous unfinished run.
The workflow must be present on the default branch for
GitHub to offer this button. See GitHub's
[manual workflow instructions](https://docs.github.com/en/actions/how-tos/manage-workflow-runs/manually-run-a-workflow).

Each run installs the documentation dependencies, refits the benchmark models,
regenerates the calibration table and the 2,000-estimator confidence-width
versus observed-error plots, and executes every Python gallery example.
Generated tables and images are build outputs, not checked-in inputs; even
local incremental builds rerun the generators and gallery scripts. The legacy
`paper/` folder is not part of this workflow.

Every selected branch publishes to the same
[documentation website](https://contrib.scikit-learn.org/forest-confidence-interval/).
To preview a PR before merging, run the action on its source branch and follow
the deployment link on the completed run. This temporarily replaces the public
documentation; it does not create a separate preview URL or merge the PR.
To restore the documentation from `master`, start a new run with `master`
selected and wait for deployment to finish. The last successfully deployed
version stays online while a build runs or if it fails. The downloadable
`github-pages` artifact is also available under **Artifacts** on each run.
Python package tests still run automatically as before.

Repository setup: GitHub Pages must use **GitHub Actions** as its source, and
**Settings → Environments → github-pages → Deployment branches and tags** must
allow the branches you want to publish (choose **No restriction** to permit
any selected branch). A `master`-only environment rule prevents PR branches
from deploying even when the workflow itself allows them.

To build locally from the repository root:

```sh
python -m pip install -e ".[docs]"
MPLBACKEND=Agg make -C docs html
```

Open `docs/_build/html/index.html` after the build. California Housing is fetched
on the first run (internet access required); scikit-learn may cache the raw
input dataset. Models, tables and plots are always recomputed. Allow several
minutes and several GiB of memory: a complete build took about 3 minutes
38 seconds on a four-core Intel MacBook Pro. The calibration benchmark alone
previously peaked near 3 GiB. GitHub runner timings may differ.
Failures in a generator or gallery example fail the build instead of publishing
incomplete results.

## Contributing

Contributions are very welcome, but we ask that contributors abide by the
[contributor covenant](http://contributor-covenant.org/version/1/4/).

To report issues with the software, please post to the
[issue log](https://github.com/scikit-learn-contrib/forest-confidence-interval/issues)
Bug reports are also appreciated, please add them to the issue log after
verifying that the issue does not already exist.
Comments on existing issues are also welcome.

Please submit improvements as pull requests against the repo after verifying
that the existing tests pass and any new code is well covered by unit tests.
Please write code that complies with the Python style guide,
[PEP8](https://www.python.org/dev/peps/pep-0008/).

E-mail [Ariel Rokem](mailto:arokem@gmail.com), [Kivan Polimis](mailto:kivan.polimis@gmail.com), or [Bryna Hazelton](mailto:brynah@phys.washington.edu ) if you have any questions, suggestions or feedback.

## Testing

The project uses a `pyproject.toml` build configuration and a `src/` package
layout. Runtime, test, documentation, and development dependencies are declared
as project dependency groups, while tests live in the top-level `tests/`
directory.

For a complete development environment, install the editable package and its
development dependencies:

```shell
python -m pip install -e ".[dev]"
python -m pytest
```

To test only against the scikit-learn version already installed in the current
environment, install the test runner and then install `forestci` without
resolving dependencies:

```shell
python -m pip install pytest==9.0.3
python -m pip install --no-deps -e .
python -m pytest
```

The GitHub Actions test matrix selects the newest bugfix release from every
scikit-learn minor series from 1.0 through 1.9. Scikit-learn 1.0–1.2 are tested
on Python 3.10 because those releases do not provide Python 3.12 wheels; all
newer series are tested on Python 3.12. A rolling `scikit-learn<2.0` job also
tests the latest available 1.x release.

## Citation

Click on the JOSS status badge for the Journal of Open Source Software article on this project.
The BibTeX citation for the JOSS article is below:

```
@article{polimisconfidence,
  title={Confidence Intervals for Random Forests in Python},
  author={Polimis, Kivan and Rokem, Ariel and Hazelton, Bryna},
  journal={Journal of Open Source Software},
  volume={2},
  number={1},
  year={2017}
}
```
