Metadata-Version: 2.4
Name: bitbullet
Version: 0.7.4
Summary: Python building blocks for ML, forecasting, clustering, evaluation, and reproducible artefacts, plus the guided BitBullet Platform.
Author-email: BitBullet <contact@bitbullet.ai>
License-Expression: MIT
Project-URL: Homepage, https://bitbullet.co.uk
Project-URL: Documentation, https://developer.bitbullet.co.uk
Project-URL: Support, https://bitbullet.co.uk/legal/contact
Keywords: machine-learning,forecasting,time-series,clustering,data-science,feature-engineering,model-evaluation,automl,no-code
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.24.0
Requires-Dist: pandas>=2.0.0
Requires-Dist: scipy>=1.11.0
Requires-Dist: scikit-learn>=1.3.0
Requires-Dist: joblib>=1.3.0
Requires-Dist: pydantic>=2.0.0
Requires-Dist: pyarrow>=12.0.0
Provides-Extra: forecast-statistical
Requires-Dist: statsforecast<3.0.0,>=2.0.0; extra == "forecast-statistical"
Provides-Extra: inference-models
Requires-Dist: xgboost>=1.7.0; extra == "inference-models"
Requires-Dist: lightgbm>=3.3.0; extra == "inference-models"
Provides-Extra: inference-cluster
Requires-Dist: hdbscan>=0.8.33; extra == "inference-cluster"
Requires-Dist: kmodes>=0.12.0; extra == "inference-cluster"
Requires-Dist: umap-learn>=0.5.3; extra == "inference-cluster"
Requires-Dist: prince>=0.13.0; extra == "inference-cluster"
Provides-Extra: train
Requires-Dist: xgboost>=1.7.0; extra == "train"
Requires-Dist: lightgbm>=3.3.0; extra == "train"
Requires-Dist: optuna>=3.0.0; extra == "train"
Requires-Dist: shap>=0.42.0; extra == "train"
Requires-Dist: tqdm>=4.65.0; extra == "train"
Requires-Dist: psutil>=5.9.0; extra == "train"
Provides-Extra: viz
Requires-Dist: matplotlib>=3.7.0; extra == "viz"
Requires-Dist: seaborn>=0.12.0; extra == "viz"
Requires-Dist: plotly>=5.14.0; extra == "viz"
Requires-Dist: ipywidgets>=8.0.0; extra == "viz"
Requires-Dist: ipython>=8.0.0; extra == "viz"
Provides-Extra: all
Requires-Dist: statsforecast<3.0.0,>=2.0.0; extra == "all"
Requires-Dist: xgboost>=1.7.0; extra == "all"
Requires-Dist: lightgbm>=3.3.0; extra == "all"
Requires-Dist: hdbscan>=0.8.33; extra == "all"
Requires-Dist: kmodes>=0.12.0; extra == "all"
Requires-Dist: umap-learn>=0.5.3; extra == "all"
Requires-Dist: prince>=0.13.0; extra == "all"
Requires-Dist: optuna>=3.0.0; extra == "all"
Requires-Dist: shap>=0.42.0; extra == "all"
Requires-Dist: tqdm>=4.65.0; extra == "all"
Requires-Dist: psutil>=5.9.0; extra == "all"
Requires-Dist: matplotlib>=3.7.0; extra == "all"
Requires-Dist: seaborn>=0.12.0; extra == "all"
Requires-Dist: plotly>=5.14.0; extra == "all"
Requires-Dist: ipywidgets>=8.0.0; extra == "all"
Requires-Dist: ipython>=8.0.0; extra == "all"
Provides-Extra: dev
Requires-Dist: pytest>=7.4.0; extra == "dev"
Requires-Dist: pytest-cov>=4.1.0; extra == "dev"
Requires-Dist: black>=23.0.0; extra == "dev"
Requires-Dist: ruff>=0.1.0; extra == "dev"
Requires-Dist: build>=1.0.0; extra == "dev"
Dynamic: license-file

# BitBullet | [bitbullet.co.uk](https://bitbullet.co.uk)

See BitBullet: [bitbullet.co.uk](https://bitbullet.co.uk)

BitBullet is a fully managed, no-code machine learning platform. Describe your
objective to the built-in AI agent, or use guided workstations, to create
classification, regression, forecasting, and clustering experiments. You stay
in control while BitBullet handles training infrastructure, tuning, diagnostics,
and portable bundles.

The **BitBullet SDK** is the open-source Python toolkit for data scientists who
want to write and orchestrate their own workflows with composable building
blocks for transformations, training, forecasting, clustering, evaluation, and
reproducible artefact metadata.

Full documentation and practical SDK tutorials are available at
[developer.bitbullet.co.uk](https://developer.bitbullet.co.uk).

## SDK Or Platform?

Use the **BitBullet SDK** when you want the raw Python building blocks to write
and orchestrate your own data-science workflows.

Use **BitBullet Platform** when you want to configure and manage a supported
modelling lifecycle through a guided, no-code environment. It centralises
projects, datasets, managed compute, storage, and repeatable workflows while
keeping you in control of every decision. Configure and compare experiments with
clicks, or ask the AI assistant to prepare a draft for review; inspect the
evidence, then export fitted artefacts, preprocessing, metadata, and generated
inference code.

[Explore BitBullet Platform](https://bitbullet.co.uk) |
[Start free](https://bitbullet.co.uk/register)

## SDK Modules

| Module                        | Purpose                                                                                                                                                         |
| ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `bitbullet.transform`       | Fitted transformation pipelines for numerical, categorical, and datetime features.                                                                              |
| `bitbullet.train`           | Supervised classification/regression training utilities, Optuna-backed search, feature selection, sample weights, threshold optimization, and training reports. |
| `bitbullet.forecast`        | Ordered and panel forecasting schemas, feature framing, reduction strategies, rolling-origin backtesting, intervals, and reconciliation.                        |
| `bitbullet.model_selection` | Ordered holdouts and expanding or rolling temporal windows with explicit gaps and audit metadata.                                                               |
| `bitbullet.cluster`         | K-Means, K-Modes, K-Prototypes, DBSCAN, GMM, gamma estimation, categorical weighting, K selection, profiling, and clustering metrics.                           |
| `bitbullet.evaluate`        | Structured classification and regression evaluation metrics for reports and metadata.                                                                           |
| `bitbullet.model`           | Model wrappers, model metadata, dataset metadata, and serialization helpers.                                                                                    |

## Installation

```bash
pip install bitbullet
```

Optional extras keep installations lean:

```bash
pip install "bitbullet[inference-models]"   # LightGBM and XGBoost wrappers
pip install "bitbullet[inference-cluster]"  # clustering extras such as kmodes
pip install "bitbullet[forecast-statistical]"  # optional StatsForecast adapter
pip install "bitbullet[train,viz]"          # training, SHAP, and plotting tools
pip install "bitbullet[all]"                # complete SDK
```

## Transform Data

```python
from bitbullet.transform import TransformPipeline

pipeline = TransformPipeline(name="credit_features")
pipeline.add("numerical", "standard_scale", columns=["income", "balance"])
pipeline.add("categorical", "onehot_encode", columns=["region"])

X_transformed = pipeline.fit_transform(X_train)
X_new = pipeline.transform(X_new_raw)
pipeline.save("artifacts/transform_pipeline.joblib")
```

Target-aware encoders receive `y` directly. `target_encode` is leakage-aware:
`fit_transform(..., y=...)` returns out-of-fold training encodings, while later
`transform(...)` calls use the stored full-training smoothed mapping.

```python
pipeline = TransformPipeline()
pipeline.add(
    "categorical",
    "target_encode",
    columns=["merchant_category"],
    params={"target_type": "classification", "cv_folds": 5, "cv_strategy": "stratified"},
)
X_encoded = pipeline.fit_transform(X_train, y=y_train)
```

## Train A Classifier

```python
from bitbullet.train import TrainConfig, OptunaTrainer

config = TrainConfig(
    name="default_risk_lgbm",
    model_type="lgbm",
    task="binary_classification",
    n_trials=30,
    optimization_metric="roc_auc",
    optuna_sampler="tpe",  # tpe, random, grid, cmaes
)

trainer = OptunaTrainer(config)
model = trainer.fit(X_train, y_train, X_val=X_val, y_val=y_val)

print(trainer.best_params)
print(trainer.state.optimal_threshold)
```

Manual fixed-parameter training is available when you do not want a search:

```python
config = TrainConfig(
    name="fixed_rf",
    model_type="random_forest",
    optimizer="manual",
    model_params={"n_estimators": 300, "max_depth": 20},
)
```

Optuna-backed grid and random search are explicit sampler choices:

```python
config = TrainConfig(
    name="small_grid",
    model_type="lgbm",
    optuna_sampler="grid",
    search_space={
        "num_leaves": [31, 63],
        "learning_rate": [0.05, 0.1],
    },
)
```

## Evaluate Classification

```python
from bitbullet.evaluate import evaluate_classification

report = evaluate_classification(
    y_true=y_test,
    y_pred_proba=model.predict_proba(X_test),
    threshold=trainer.state.optimal_threshold or 0.5,
)

metadata_ready = report.to_dict()
```

## Train And Evaluate A Regressor

```python
from bitbullet.evaluate import evaluate_regression
from bitbullet.train import TrainConfig, OptunaTrainer

config = TrainConfig(
    name="house_price_lgbm",
    model_type="lgbm",
    task="regression",
    n_trials=30,
    optimization_metric="rmse",  # minimize by default for regression
    optuna_sampler="tpe",
)

trainer = OptunaTrainer(config)
model = trainer.fit(X_train, y_train)

y_pred = model.predict(X_test)
report = evaluate_regression(
    y_true=y_test,
    y_pred=y_pred,
    n_features=X_train.shape[1],
)

metadata_ready = report.to_dict()
```

## Forecast Ordered And Panel Data

```python
from sklearn.linear_model import Ridge

from bitbullet.forecast import (
    ForecastConfig,
    ForecastFeatureBuilder,
    ForecastFrame,
    ForecastSchema,
    RollingFeature,
    TabularForecaster,
)

schema = ForecastSchema(
    time="date",
    targets="sales",
    entities="store",
    future_covariates=("promotion", "temperature"),
    cadence="D",
)
history = ForecastFrame(history_df, schema)
features = ForecastFeatureBuilder(
    target_lags=(1, 7),
    rolling_features=(RollingFeature("sales", window=7, lag=1),),
    calendar_features=("day_of_week_sin", "day_of_week_cos"),
)

forecaster = TabularForecaster(
    Ridge(alpha=1.0),
    config=ForecastConfig(horizons=(1, 2, 3), strategy="direct"),
    feature_builder=features,
).fit(history)

predictions = forecaster.forecast(future_df)
```

`ForecastBacktester` adds synchronized rolling-origin evaluation, while
`ConformalIntervalCalibrator` and `HierarchicalReconciler` provide optional
interval and aggregation layers. See the
[forecasting API guide](https://developer.bitbullet.co.uk/api/forecast/) and
advanced forecasting tutorial for the complete workflow.

## Cluster Data

```python
from bitbullet.cluster.core import ClusterConfig
from bitbullet.cluster.algorithms.partitional import KPrototypesClusterer

config = ClusterConfig(
    name="customer_segments",
    algorithm_type="partitional",
    method="kprototypes",
    n_clusters=5,
    numerical_columns=["income", "spend"],
    categorical_columns=["region", "channel"],
    params={
        "gamma": "huang",
        "categorical_weights": "relevance",
        "init": "Cao",
        "n_init": 10,
    },
)

clusterer = KPrototypesClusterer(config)
labels = clusterer.fit_predict(df)

print(clusterer.state.fitted_params["gamma_by_column"])
print(clusterer.state.fitted_params["categorical_weights_by_column"])
```

## Save Models With Metadata

```python
from bitbullet.model import ModelMetadata, ModelSerializer

metadata = ModelMetadata(
    name="default_risk_lgbm",
    model_type=model.model_type,
    framework=model.framework,
    task="binary_classification",
    metrics=report.metrics,
)
metadata.add_feature_schema(X_train)

ModelSerializer.save(
    model=model,
    path="artifacts/default_risk_lgbm.pkl",
    metadata=metadata,
    train_data=(X_train, y_train),
    test_data=(X_test, y_test),
    include_datasets=False,
)
```

## License

MIT

## Support

Questions and problem reports can be sent to contact@bitbullet.ai.
