Metadata-Version: 2.4
Name: hessboost
Version: 0.2.0
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: Free Threading :: 3 - Stable
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Programming Language :: Rust
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Dist: numpy>=2
Requires-Dist: pandas>=2.2 ; extra == 'pandas'
Requires-Dist: scikit-learn>=1.6 ; extra == 'scikit-learn'
Provides-Extra: pandas
Provides-Extra: scikit-learn
Summary: Fast, deterministic gradient boosting in Rust: conformal intervals, explainable boosting machines, distributional boosting, and tree-based diffusion
Keywords: xgboost,gradient-boosting,gbdt,machine-learning,decision-trees,rust
Author-email: Brenden Matthews <brenden@brndn.io>
License-Expression: Apache-2.0
Requires-Python: >=3.11
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Changelog, https://github.com/brndnmtthws/hessboost/releases
Project-URL: Documentation, https://github.com/brndnmtthws/hessboost/tree/main/python#readme
Project-URL: Homepage, https://github.com/brndnmtthws/hessboost
Project-URL: Issues, https://github.com/brndnmtthws/hessboost/issues
Project-URL: Repository, https://github.com/brndnmtthws/hessboost

# hessboost for Python

**Fast, deterministic gradient boosting in Rust.** The core API (`DMatrix`,
`train`, `cv`, `Booster`, and scikit-learn estimators) takes the parameter
names XGBoost users know, and models move to and from XGBoost as JSON or
UBJSON.

- **Strict.** An unknown parameter, a value of the wrong type or range, or
  a combination hessboost does not implement raises an error; nothing is
  silently ignored.
- **Deterministic.** The same parameters, data, and seed give the same
  model at any `nthread`.
- **Typed** (`py.typed`, complete type information), with the GIL released
  while training and predicting, and free-threaded CPython supported.
- **Modern modeling (opt-in).** Conformal prediction intervals,
  confidence intervals for the regression function (Boulevard boosting),
  distributional boosting (a predictive distribution per row), LightGBM/CatBoost
  tree options, class-balanced binary bagging, and XE-NDCG ranking
  (`objective="rank:xendcg"`).

## Installation

```sh
uv add hessboost        # or: pip install hessboost
```

Extras: `hessboost[pandas]` and `hessboost[scikit-learn]`. numpy is the only
required dependency.

Prebuilt wheels are published for Linux x86_64 and aarch64 (glibc
manylinux and musl/Alpine musllinux), macOS arm64 (with Metal support,
`device="metal"`), and Windows x86_64: one `abi3` wheel per platform for
CPython 3.11 and newer, plus a wheel for free-threaded CPython 3.14t.
Elsewhere the installer builds from the source distribution, which needs
Rust 1.93 or newer and a C compiler (for libzstd).

## Quick start

```python
import numpy as np
import hessboost

rng = np.random.default_rng(0)
X = rng.normal(size=(1000, 5))
y = (X[:, 0] + X[:, 1] ** 2 > 1).astype(int)

dtrain = hessboost.DMatrix(X[:800], label=y[:800])
dvalid = hessboost.DMatrix(X[800:], label=y[800:])

booster = hessboost.train(
    {"objective": "binary:logistic", "max_depth": 4, "eta": 0.1, "eval_metric": "auc"},
    dtrain,
    num_boost_round=500,
    evals=[(dvalid, "valid")],
    early_stopping_rounds=20,
    verbose_eval=False,
)
probabilities = booster.predict(X[800:])  # through booster.best_iteration
shap = booster.predict(X[800:], pred_contribs=True)  # (rows, features + 1)
print(booster.best_iteration, booster.get_score(importance_type="gain"))
```

`DMatrix` takes numpy arrays of any numeric dtype and memory layout (a
C-contiguous `float32` array is used without a copy), pandas DataFrames,
scipy sparse matrices, and anything `numpy.asarray` accepts. NaN is missing
(or pass `missing=`); ranking data takes `group=` sizes or `qid=`, and
`survival:aft` takes `label_lower_bound=`/`label_upper_bound=`.
`Booster.predict` accepts the same inputs directly.

`train` supports XGBoost's everyday arguments: `evals`, `evals_result`,
`early_stopping_rounds`, `verbose_eval` (printed live, every round or
every `n`-th), `xgb_model` (continued training, or tree refresh with
`process_type="update"`), a custom objective `obj`, a `custom_metric`,
and `callbacks`. Ctrl-C stops training at the end of the current round
and raises `KeyboardInterrupt`. `hessboost.cv` cross-validates over
shuffled folds, explicit folds, or a scikit-learn splitter.

```python
class StopAtTarget(hessboost.TrainingCallback):
    def after_iteration(self, iteration, evals_log):
        return evals_log["valid"]["auc"][-1] > 0.99  # True stops training


hessboost.train(params, dtrain, 1000, evals=[(dvalid, "valid")], callbacks=[StopAtTarget()])
```

### pandas and categorical features

DataFrame column names become feature names, and `category` columns become
native categorical features. A category's code is its position in the
column's categories, so the booster remembers each column's categories and
re-codes a frame passed to `predict` (and the other prediction and
calibration methods) whose categories are ordered differently or include
unseen values (which count as missing):

```python
import pandas as pd

df = pd.DataFrame({"color": pd.Categorical(["red", "blue", "red"] * 100), "size": np.arange(300.0)})
booster = hessboost.train({}, hessboost.DMatrix(df, label=np.arange(300.0)), 20)
booster.predict(df)
```

A `DMatrix` is coded once, when it is built, so it must have the features
of every model or matrix it meets: `predict` and the calibrators check it
against each model reading it, and `train` checks every `evals` matrix
against `dtrain` and against `xgb_model`, and `dtrain` against `xgb_model`
(continued training or refresh). Feature names (where both have them;
`predict(validate_features=False)` skips them), which features are
categorical, and each categorical feature's categories, in order, must
match; a mismatch raises `HessboostError` naming the eval set and the
feature. Build eval frames with the training frame's categories (for
example `valid["color"].cat.set_categories(train["color"].cat.categories)`).
Codes without recorded categories (numpy data with
`feature_types=["c", ...]`) are taken to be the other side's codes; only
which features are categorical is compared, and a model continued on them
keeps the earlier model's categories. `ConformalizedQuantile.calibrate`
likewise needs its two models to share their features, and re-codes frames
to whichever model records categories. The scikit-learn estimators re-code
every `eval_set` frame to the training frame's categories and, with
`xgb_model`, the training frame to the earlier model's.

## scikit-learn

```python
from hessboost.sklearn import HessboostClassifier

model = HessboostClassifier(
    n_estimators=300, max_depth=4, learning_rate=0.1, early_stopping_rounds=20
)
model.fit(X_train, y_train, eval_set=[(X_valid, y_valid)])
model.predict_proba(X_test)
model.feature_importances_
```

`HessboostRegressor` (multi-target with a 2-D `y`), `HessboostClassifier`
(any class labels), `HessboostRanker` (`group=` or `qid=`), and
`HessboostDistributionRegressor` (`dist:*` objectives; `predict` returns
means, `predict_distribution` the distributions) take XGBoost's
scikit-learn parameter names, plus `callbacks`. Parameters left at `None`
keep hessboost's defaults, and `params={...}` passes any other training
parameter; `fit(..., verbose=True)` (or a period) prints rounds live. They
work with pipelines, `clone`, grid search, and pickling, and pass
scikit-learn's estimator checks (except three documented deviations).
`import hessboost` does not import scikit-learn; `hessboost.sklearn` needs
it.

## Model files

| Method | Formats |
|---|---|
| `save_model(path, format=None)` | by extension: `.json` native JSON, `.ubj` XGBoost UBJSON, anything else native binary; or `format="binary" \| "json" \| "xgboost-json" \| "xgboost-ubjson"` |
| `Booster(path_or_bytes)`, `load_model(...)` | any of the four, or a LightGBM 4.x text model (`format="lightgbm"`, import only), detected from the content (`ModelFormatError` if it looks like none of them) |
| `save_raw(format="binary")` | the same formats as bytes |
| `pickle` / `copy` | native binary plus feature names, categories, and `best_score` |

The native binary format is compressed, checksummed, and lossless; files
load in subsequent releases. Use
`format="xgboost-json"` or `.ubj` for a file XGBoost loads. A LightGBM
model (`lightgbm.Booster.save_model`) predicts LightGBM's values for
missing values as `NaN` and categorical features as non-negative codes;
models with no exact equivalent raise `ModelFormatError`. `booster[a:b]`
slices boosting iterations.

## Modern modeling

```python
from hessboost.conformal import ConformalizedQuantile

band = hessboost.train(
    {"objective": "reg:quantileerror", "quantile_alpha": [0.05, 0.95]},
    hessboost.DMatrix(X_train, y_train),
    200,
)
cqr = ConformalizedQuantile.calibrate_outputs(band, X_cal, y_cal, alpha=0.1)
lower, upper = cqr.predict_interval(X_test).T  # >= 90% coverage, finite-sample
lower, upper = inference.confidence_intervals(X_test, alpha=0.05).T  # for f(x)

dist = hessboost.train({"objective": "dist:normal"}, hessboost.DMatrix(X_train, y_train), 300)
d = dist.predict_distribution(X_test)
d.mean(), d.std(), d.interval(0.9), d.log_prob(y_test), d.crps(y_test)

sglb = hessboost.train({"posterior_sampling": True}, hessboost.DMatrix(X_train, y_train), 1000)
members, iterations = sglb.predict_virtual_ensembles(X_test, 10)  # (10, rows)
u = sglb.predict_uncertainty(X_test, 10)  # u.knowledge rises off the training data

from hessboost.online import Approximate, OnlineModel

online = OnlineModel.train(
    {"tree_method": "hist", "max_depth": 6},
    hessboost.DMatrix(X_train, y_train),
    100,
    Approximate(0.1),
)
report = online.update(hessboost.DMatrix(X_new, y_new), deletions=[3, 17])
online.model.predict(X_test)  # online.data: the updated training rows

from hessboost.inference import BoulevardInference, honest_refit

params = {"booster": "boulevard", "eta": 0.8, "boulevard_dropout": 0.5, "subsample": 0.8}
trained = hessboost.train(params, hessboost.DMatrix(X_struct, y_struct), 200)
model = honest_refit(trained, X_values, y_values)  # leaves from independent rows
inference = BoulevardInference.fit(model, X_values, holdout=X_cal, holdout_label=y_cal)

from hessboost.diffusion import DiffusionModel, DiffusionParams, crps, quantiles

flow = DiffusionModel.fit(DiffusionParams.flow_matching(), X_train, y_train)
draws = flow.sample(X_test, 200, seed=0)  # (rows, 200, outputs) float32
quantiles(draws, [0.05, 0.5, 0.95])  # (rows, 3, outputs)
crps(draws, y_test)  # (rows, outputs)

from hessboost.diffusion.forest import ForestModel, ForestParams

forest = ForestModel.fit(ForestParams.forest_diffusion(), X_with_nans)
synthetic = forest.sample(1000, seed=0)  # ForestSamples: .values (1000, columns), .labels None
filled = forest.impute(X_with_nans, n_imputations=5)  # (5, rows, columns)
```

- `hessboost.conformal`: `SplitConformal` and `ConformalizedQuantile`
  (from two quantile models, two outputs of one, or a `dist:*` model).
- `hessboost.online`: `OnlineModel` adds and deletes training rows of a
  trained model in place (incremental learning, machine unlearning).
  `mode=Exact()` is exact: every update equals `hessboost.train` on
  `online.data` bit for bit; `Approximate(tolerance)` (the default, at
  `0.1`) keeps splits that still rank near the top and is faster than
  retraining for small changes. `update(additions, deletions, callback=...)`
  returns an `UpdateReport` (`nodes_kept`, `subtrees_regrown`,
  `rows_refreshed`), or
  `None` when `callback(iteration)` returned `True`; a refused, stopped, or
  interrupted (Ctrl-C) update changes nothing. `OnlineModel.from_model`
  resumes from a saved `Booster` and its training data. Updates need `hist`
  depth-wise trees without sampling or constraints and unweighted data.
- `hessboost.ebm`: explainable boosting machines (`{"booster": "ebm"}`,
  cyclic GA2M with outer bags, FAST pairs, per-bag early stopping, and
  categorical terms): `shape_functions` returns every term's
  piecewise-constant shape (`NumericAxis` edges or `CategoricalAxis`
  codes, plus a missing cell per axis) and `Booster.ebm`;
  `hessboost.inference.EbmInference` puts confidence bands on the shapes of
  an `ebm_boulevard` model.
- `hessboost.inference`: Boulevard boosting's asymptotic confidence,
  prediction (Gaussian noise), and reproduction intervals for `f(x)`
  (`BoulevardInference`, exact or Nystrom), `honest_refit`, and
  `Booster.boulevard`. The intervals are conditional on the tree
  structures: nominal for low-dimensional smooth signals after an honest
  refit, under-covering elsewhere (see the crate's `inference` docs).
- `hessboost.diffusion`: nonparametric `p(y | x)` for scalar or vector
  labels (multimodal, skewed, heavy-tailed) by conditional diffusion or
  flow matching with GBDT score models, after Treeffuser and DiffGBM.
  `DiffusionParams` is a frozen dataclass tree (`Score`/`FlowMatching` and
  their SDEs, paths, and time sampling; `EarlyStopping`, `Residualizer`)
  with presets `default()`, `treeffuser()`, and `flow_matching()`; its
  `training` mappings are XGBoost parameters, as `train` reads them.
  `DiffusionModel.sample` returns `(rows, n_samples, outputs)` draws,
  deterministic per seed and step count (`n_steps=` overrides the model's);
  `mean`, `quantiles`, and `crps` summarize them.
  Models save with `to_bytes(format="binary")` / `save(path,
  format="binary")` (`"binary"` or `"json"`) and load with
  `from_bytes(data, format="auto")` / `load(path, format="auto")`, which
  detect the format (`ModelFormatError` for bytes in neither), or pickle
  (which also keeps feature names and categories). `fit` releases
  the GIL but cannot be interrupted: Ctrl-C takes effect once it returns.
- `hessboost.diffusion.forest`: ForestFlow / ForestDiffusion synthetic
  tabular rows (optionally per class of a label) and missing-value
  imputation with per-noise-level GBDTs. `ForestParams` (method `"flow"`
  or `Diffusion(beta_min, beta_max)`, `n_t`, `duplicate_k`, one
  `"continuous"`/`"integer"`/`"categorical"` kind per column, XGBoost
  `training` parameters) has presets `forest_flow()` (the defaults) and
  `forest_diffusion()`. `ForestModel.sample(n_rows)` returns a frozen
  `ForestSamples(values, labels)` (labels only for a class-conditional
  model), `sample_for_labels(labels)` one row per label, and `impute(X, y, n_imputations=, repaint=Repaint(...))`
  `(n_imputations, rows, columns)` with the observed entries kept
  (diffusion only). Values are in the data's own coding. Models save and
  load like `DiffusionModel`s.
- `hessboost.folds`: `k_fold`, `forward_chaining` (expanding-window,
  purged by a row `gap`), and `purged_forward` (timestamped rows, purged
  by each row's own label window, for overlapping or irregular horizons)
  folds for `cv` or your own validation loops.
- `Booster.predict_virtual_ensembles` / `predict_uncertainty`: CatBoost's
  virtual ensembles of an SGLB model (`posterior_sampling`, `langevin`,
  `model_shrink_rate`), with knowledge, data, and total uncertainty
  (`hessboost.Uncertainty`).
- Every hessboost training option (`path_smooth`, `extra_trees`,
  `linear_tree`, `grow_policy="symmetric"`, `use_quantized_grad`,
  `pos_bagging_fraction`, `neg_bagging_fraction`, `bagging_by_query`, the
  `dist:*` objectives and their `dist_gradient`, `rank:xendcg`, ...) is a
  `params` key. XE-NDCG's keyed per-round random stream differs from
  LightGBM's `rank_xendcg` stream.

## Differences from XGBoost's Python package

- Errors: `hessboost.HessboostError` (a `ValueError`) for refused inputs,
  with subclasses `InvalidDataError` (data content: labels outside the
  objective's domain, negative weights, bad groups or bounds; the message
  names the input and any eval set), `IncompatibleModelError` (a model the
  parameters or data do not match: continued training, refresh, slicing)
  and `ModelFormatError` (model files); invalid or conflicting parameters
  raise `HessboostError` itself. `TypeError` for wrong types.
  Unsupported parameters are refused, including `verbosity`; `missing`
  belongs to `DMatrix`.
- `TrainingCallback.after_iteration(iteration, evals_log) -> bool` sees the
  round and the evaluation history, not the model (XGBoost's also gets the
  booster and has `before_training`/`after_training` hooks); returning
  `True` stops training with the rounds so far. `hessboost.cv` has no
  callbacks and finishes its folds before Ctrl-C takes effect.
- `custom_metric(predictions, labels, weights)` returns a float and is named
  by the function's `__name__` (XGBoost passes a `DMatrix` and returns
  `(name, value)`). `obj(margins, dtrain)` matches XGBoost; the model then
  predicts margins from a zero intercept (or `base_score`). `obj` replaces
  `objective`, which `params` must then not set, and `num_class` is the
  custom objective's output count (XGBoost's custom-softmax convention;
  default: one per label column).
- `cv` returns a dict of numpy arrays (`test-<metric>-mean`/`-std`) with
  held-out metrics only; there is no `stratified` or `as_pandas`.
- `predict` defaults to the iterations through `best_iteration` (XGBoost's
  scikit-learn behavior); pass `iteration_range=(0, 0)` for all.
  `pred_leaf` returns `int32`.
- Model files do not store feature names or categories (pickles do).
- Not available: `DMatrix` from files or `QuantileDMatrix`, `inplace_predict`
  (`predict` takes arrays directly), `Booster.get_dump`/`trees_to_dataframe`
  /`dump_model`, attributes (`set_attr`), plotting, distributed (Dask/Spark)
  and GPU (CUDA) training, `approx_contribs`, and `strict_shape`.

## Development

From `python/` in the [repository](https://github.com/brndnmtthws/hessboost),
with [uv](https://docs.astral.sh/uv/):

```sh
uv sync                        # build the extension and install dev tools
uv run pytest
uv run pyright --verifytypes hessboost --ignoreexternal
```

Lint, format, and type-check all of the repository's Python from its root
(configuration: `ruff.toml`, `ty.toml`):

```sh
uv run --project python ruff check
uv run --project python ruff format --check
uv run --project python ty check
```

## License

Apache-2.0.

