Metadata-Version: 2.4
Name: pandas-booster
Version: 0.2.0
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Rust
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Dist: pandas>=1.5.0,<3.0.0
Requires-Dist: numpy>=1.21.0,<3.0.0
Requires-Dist: tomli>=2.0 ; python_full_version < '3.11'
Requires-Dist: polars>=0.20 ; extra == 'bench'
Requires-Dist: pyarrow>=12.0 ; extra == 'bench'
Requires-Dist: matplotlib>=3.5 ; extra == 'bench'
Requires-Dist: pytest>=7.0 ; extra == 'dev'
Requires-Dist: pytest-benchmark>=4.0 ; extra == 'dev'
Requires-Dist: polars>=0.20 ; extra == 'dev'
Requires-Dist: basedpyright>=1.37.1 ; extra == 'dev'
Requires-Dist: pandas-stubs>=2.0.0 ; extra == 'dev'
Requires-Dist: ruff>=0.9.0 ; extra == 'dev'
Provides-Extra: bench
Provides-Extra: dev
License-File: LICENSE
Summary: High-performance numerical acceleration library for Pandas using Rust
Keywords: pandas,numpy,performance,rust,data-science,acceleration
License: MIT
Requires-Python: >=3.9
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Homepage, https://github.com/snowykr/pandas-booster
Project-URL: Issues, https://github.com/snowykr/pandas-booster/issues
Project-URL: Repository, https://github.com/snowykr/pandas-booster

# pandas-booster

[![CI](https://github.com/snowykr/pandas-booster/actions/workflows/ci.yml/badge.svg)](https://github.com/snowykr/pandas-booster/actions)
[![Python Versions](https://img.shields.io/pypi/pyversions/pandas-booster.svg)](https://pypi.org/project/pandas-booster/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

pandas-booster is a high-performance numerical acceleration library for Pandas that offloads heavy computations to Rust. It leverages multi-core parallelism and zero-copy data access to provide significant speedups for large-scale data processing tasks.

This project is an independent third-party package and is not affiliated with, endorsed by, or sponsored by the pandas project or NumFOCUS.

## Features

- Parallel GroupBy aggregations using Rayon (single and multi-column)
- Fast hashing with AHash
- Zero-copy interop between NumPy and Rust
- Release of the Python Global Interpreter Lock (GIL) during computation
- Seamless integration as a Pandas DataFrame accessor

## Installation

### From PyPI
Install the latest published release from PyPI:
```bash
pip install pandas-booster
```

If your project uses uv:
```bash
uv add pandas-booster
```

For a uv-managed environment without adding a project dependency:
```bash
uv pip install pandas-booster
```

### Development Setup
To build and install from source, **all development commands in this repository assume you are using an activated virtual environment** (I recommend `.venv`).

```bash
# 1. Create and activate virtual environment
python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate

# 2. Install build tools and dependencies
pip install "maturin>=1.13,<2.0"
pip install -e ".[bench,dev]"

# 3. Build and install in development mode
maturin develop --release
```

Equivalent setup with uv:
```bash
# 1. Create/sync the project environment with development extras
uv sync --extra bench --extra dev

# 2. Build and install the Rust extension in development mode
uv run --with "maturin>=1.13,<2.0" maturin develop --release
```

Use release builds for packaging, validation, and before comparing runtime
performance:
```bash
uv run --with "maturin>=1.13,<2.0" maturin build --release --out dist
cargo build --release --features extension-module
```

## Quick Start

```python
import pandas as pd
import numpy as np
import pandas_booster

# Create a large dataset
n = 1_000_000
df = pd.DataFrame({
    "key": np.random.randint(0, 1000, size=n),
    "value": np.random.random(size=n)
})

# Use the booster accessor for accelerated groupby
result = df.booster.groupby(by="key", target="value", agg="sum")

print(result)
```

### Multi-Column GroupBy

```python
# Create dataset with multiple key columns
n = 1_000_000
df = pd.DataFrame({
    "region": np.random.randint(0, 50, size=n),
    "category": np.random.randint(0, 100, size=n),
    "year": np.random.randint(2020, 2025, size=n),
    "sales": np.random.random(size=n) * 1000
})

# Group by multiple columns - returns Series with MultiIndex
result = df.booster.groupby(by=["region", "category"], target="sales", agg="sum")

print(result.head())
# region  category
# 0       0           9823.45
#         1           10234.12
#         2           9567.89
# ...

# Access specific groups
print(result.loc[(0, 1)])  # Sales for region=0, category=1
```

## API Reference

### `df.booster.groupby(by, target, agg, sort=True)`

Performs a Rust-accelerated groupby aggregation.

| Parameter | Type | Description |
|-----------|------|-------------|
| `by` | `str \| list[str]` | Column name(s) to group by. All columns must be integer dtype. |
| `target` | `str` | Name of the column to aggregate. Must be numeric (int or float). |
| `agg` | `str` | Aggregation function name. |
| `sort` | `bool` | If `True` (default), sort result by group keys. If `False`, preserve Pandas appearance order (first-seen group order). |

### Configuration

`pandas-booster` defaults to Rust-side sorting kernels for `sort=True`.

Emergency toggle (panic button):

- `PANDAS_BOOSTER_FORCE_PANDAS_SORT` (default: OFF when unset):
  - truthy (`1/true/yes/on`, case-insensitive): force Python `Series.sort_index()` after Rust aggregation for `sort=True`.
  - anything else (including `0/false/no/off`): keep Rust-side sorting.
- `PANDAS_BOOSTER_FORCE_PANDAS_FLOAT_GROUPBY` (default: OFF when unset):
  - truthy (`1/true/yes/on`, case-insensitive): force pandas fallback for **single-key float** `sum`/`mean`/`prod`/`std`/`var`/`median`.
  - anything else (including `0/false/no/off`): use Rust deterministic reduction kernels for those eligible operations.

Single-key float `prod` uses a verified Rust path that preserves per-key row-order IEEE-754 overflow/underflow semantics without merging partial products; multi-key float `prod` remains Rust-eligible. These toggles are intended for quick rollback if a Rust ordering or deterministic-float reduction issue is discovered. Forcing Python sort moves the `sort=True` cost to Pandas and is slower. Forcing pandas float groupby rolls single-key float `sum`/`mean`/`prod`/`std`/`var`/`median` back to pandas semantics/performance.

Note: the Rust-side `sort=True` kernels allocate a permutation vector and perform an `O(G log G)` comparison sort over groups (G = number of groups). This can increase memory usage at very high cardinality.

Note: the benchmark runner defaults `PANDAS_BOOSTER_FORCE_PANDAS_SORT=0`.

ABI skew controls:

- `PANDAS_BOOSTER_STRICT_ABI` (default: OFF when unset):
  - truthy (`1/true/yes/on`, case-insensitive): treat detected ABI skew as a hard error (no fallback).
  - anything else (including `0/false/no/off`): fall back to pandas on detected ABI skew.
- `PANDAS_BOOSTER_ABI_SKEW_NOTICE` (default: ON when unset):
  - unset / truthy (`1/true/yes/on`, case-insensitive): enable ABI-skew warnings.
  - anything else (including `0/false/no/off`): disable ABI-skew warnings.

**Returns**: 
- Single key (`by="col"`): A `pd.Series` indexed by the unique keys.
- Multiple keys (`by=["col1", "col2"]`): A `pd.Series` with a `pd.MultiIndex`.

## Supported Operations

The following aggregation functions are currently supported:

| Operation | Description |
|-----------|-------------|
| `sum` | Sum of values in each group |
| `mean` | Arithmetic mean of values in each group |
| `median` | Median of values in each group |
| `prod` | Product of values in each group |
| `std` | Sample standard deviation (ddof=1) |
| `var` | Sample variance (ddof=1) |
| `min` | Minimum value in each group |
| `max` | Maximum value in each group |
| `count` | Count of non-NaN values in each group |

### Acceleration Scope and Mandatory Fallback

To ensure predictable performance and correctness, the following rules define where Rust-first dispatch is certified.

#### Certified Rust Dispatch Domain (`std`/`var`/`median` and single-key float `prod`)

| Feature | Supported (Rust-first) | Mandatory pandas Fallback |
|---------|-------------------------|--------------------------|
| Keys | Single or Multi-key (up to 10), integer dtypes | Non-integer keys, custom objects |
| Values | Numeric (`int64`, `float64`) | `uint64`, object, bool, datetime, category |
| Semantics | pandas defaults (`std`/`var`: `ddof=1`; `median`: skip `NaN` values) | Custom `ddof`, `numeric_only`, `skipna=False` |
| Dtypes | Primitive NumPy arrays | Extension dtypes (nullable `Int64`, `Float64`, `pd.NA`) |

#### Determinism and Precision

- **Determinism**: Accelerated float aggregations (`sum`, `mean`, `prod`, `std`, `var`, `median`) are bitwise-identical across thread counts for identical inputs in the same environment. Single-key float `prod` is accelerated with a no-merge path that preserves pandas row-order product semantics per key.
- **Precision**: Results are semantically pandas-compatible but not guaranteed to be bit-for-bit identical to pandas. `std` and `var` follow this same policy.
- **Escape Hatch**: `PANDAS_BOOSTER_FORCE_PANDAS_FLOAT_GROUPBY=1` forces pandas execution for single-key float-input `sum`/`mean`/`prod`/`std`/`var`/`median` if bit-for-bit identity is required.

## Requirements and Constraints

To ensure correctness and performance, the following constraints apply:

- **Minimum dataset size**: 100,000 rows for legacy aggregations (`sum`, `mean`, `prod`, `min`, `max`, `count`). For smaller datasets on these operations, the library automatically falls back to native Pandas. **Note**: Supported `std`, `var`, and `median` operations are Rust-first by default within their certified dispatch domain regardless of dataset size.
- **Key column(s)**: Must be integer dtype (e.g., `int64`, `int32`). For multi-column groupby, all key columns must be integers. The accelerated path preserves Pandas' index dtype (e.g., `int32` on numpy-backend pandas).
- **Maximum key columns**: Up to 10 columns for multi-column groupby.
- **Value column**: Must be a numeric dtype (integers or floats).
- **Extension dtypes**: Pandas extension dtypes (e.g., nullable `Int64` / `Float64` using `pd.NA`) are not supported and will trigger a fallback to Pandas.
- **NaN handling**: `NaN` values in the target column are skipped in aggregations, matching standard Pandas behavior.
- **Determinism policy (single-key float `sum`/`mean`/`prod`/`std`/`var`/`median`)**: For identical inputs in the same runtime environment, pandas-booster returns bitwise-identical results across thread counts. `NaN` inputs are skipped; all-`NaN` groups follow existing semantics (`sum -> +0.0`, `prod -> 1.0`, `mean -> NaN`, `std/var/median -> NaN`). Compared with pandas, accelerated `sum`/`mean`/`std`/`var` outputs may differ at the last-bit level (including `+0.0` vs `-0.0`) because pandas-booster uses an implementation-defined deterministic reduction order; single-key float `prod` preserves pandas row-order product semantics per key.
- **Return types**: Integer aggregations follow Pandas-style dtypes: `sum/prod/min/max/count` return integer results for signed integer inputs; all unsigned `prod` inputs always fall back to pandas; `mean/std/var/median` return `float64`. Standard deviation and variance always use `ddof=1`.

## Performance

The library is designed for large datasets where multi-core parallelism can be fully utilized.

- `sort=True`: single-key groupby uses Rayon's parallel map-reduce; multi-key groupby uses a radix-partitioning algorithm that eliminates merge overhead.
- `sort=False`: results preserve Pandas appearance order (first-seen group order). Internally, this path tracks the first-seen row index per group and reorders groups with an integer radix sort to avoid `O(G log G)` comparison sorting.

**Benchmark methodology:**
- **Process Isolation:** Benchmarks use rigorous process isolation to ensure accurate results.
- **Host machine:** MacBook Pro (Mac15,6), Apple M3 Pro, 11 CPU cores (5 Performance + 6 Efficiency), 18 GB RAM, macOS 26.4.1.
- **Samples:** The checked-in benchmark reports are publication-quality local artifacts generated with `--samples 20 --cardinality all --sort-mode all`. Each sample runs in a fresh Python process.
- **Cold:** Average of the requested fresh process executions (1st run measured immediately).
- **Warm:** Average of the requested fresh process executions. Each process runs Cold once and a Warmup once (both discarded), then measures the next run (steady state).
- **Correctness:** Booster and Polars outputs are validated against a Pandas baseline. For `sort=False`, benchmarks validate Pandas-compatible appearance order (first-seen group order).
- **Polars sort handling:** Polars does not have a `sort` parameter in `group_by`. For fair comparison, I define `sort=True` as "groupby+agg followed by sorting the result by keys" (cost included in timing), and `sort=False` as "groupby+agg with Pandas-compatible appearance order (first-seen group order)". This ensures all three engines (Pandas, Polars, Booster) are measured under identical conditions.
- **Profile evidence:** `--profile-json` writes internal single-key `std`/`var` phase timings for the Rust path (`local_build`, `merge`, `reorder`, `materialize`, and Python post-processing) so benchmark reports can separate kernel time from conversion and Series construction overhead.
- **Speedup baseline:** All speedup values (`x`) use **Pandas** as the baseline (1.0x) within each sort mode.
- **Optional Polars:** Polars is included in the benchmarks for comparison if installed. If not installed, the benchmark suite proceeds with Pandas vs Booster only.

Benchmark tables are stored as per-aggregation reports under [`benchmarks/reports/`](benchmarks/reports/).
Generated benchmark Markdown files carry a provenance marker, and reruns only overwrite or delete
files with that marker so custom output directories cannot silently clobber unrelated `README.md`
or `<agg>.md` files.

### Sorted vs Appearance-Ordered Results

By default, results are sorted by group keys to match Pandas `sort=True` output. Pass `sort=False` to preserve Pandas' appearance order (first-seen group order). Performance impact depends on workload (see the linked benchmark reports above):

```python
# Sorted (default) - matches Pandas exactly
result = df.booster.groupby(by=["a", "b"], target="val", agg="sum")

# Appearance-ordered (sort=False) - matches Pandas semantics
result = df.booster.groupby(by=["a", "b"], target="val", agg="sum", sort=False)
```

### Benchmark Reproduction

To reproduce the checked-in benchmark reports:

```bash
# Install benchmark dependencies and build in release mode
pip install -e ".[bench,dev]"
maturin develop --release

# Or, with uv
uv sync --extra bench --extra dev
uv run --with "maturin>=1.13,<2.0" maturin develop --release

# Run the checked-in publication-quality reports for all supported aggregations
python benchmarks/generate_docs.py --samples 20 --cardinality all --sort-mode all

# Run lightweight smoke reports when iterating locally
python benchmarks/generate_docs.py --samples 1 --cardinality standard --sort-mode sorted

# Run default sum benchmark only (standard + high)
python benchmarks/benchmark.py --samples 20 --output benchmarks/reports

# Run only selected aggregation functions
python benchmarks/benchmark.py --agg std --agg var --samples 20 --output benchmarks/reports
python benchmarks/benchmark.py --agg median --samples 20 --output benchmarks/reports

# Save single-key std/var phase-profile evidence as JSON
python benchmarks/benchmark.py --agg std --agg var --samples 20 --profile-json profile.json

# Include threshold diagnostics as well
python benchmarks/benchmark.py --cardinality all --diagnostic threshold --sort-mode unsorted --samples 20 --output benchmarks/reports
```

#### Environment & Configuration
The following environment was used to generate the checked-in benchmark reports.
(Note: These are **not** the minimum requirements for using the library, but strictly the environment used for reproduction).

- **Build Mode**: Release (`maturin develop --release`)
- **Machine**: MacBook Pro (`Mac15,6`), Apple M3 Pro, 11 CPU cores (5 Performance + 6 Efficiency), 18 GB RAM
- **Threading**: Default Rayon behavior (uses all available logical cores)
- **OS**: macOS 26.4.1 (Darwin 25.4.0, arm64)
- **Python**: 3.11.15
- **Pandas**: 2.3.3
- **Polars**: 1.40.1

**Note**: The benchmark scripts use rigorous process isolation (fresh process per sample) for both Cold and Warm measurements to ensure accurate results.

For more detailed benchmark options and configurations, see the [Development > Benchmarking](#benchmarking) section.

## Development

### Building
Build the extension module in-place:
```bash
source .venv/bin/activate
maturin develop

# Or, with uv
uv run --with "maturin>=1.13,<2.0" maturin develop
```

For release builds with optimizations:
```bash
source .venv/bin/activate
maturin develop --release

# Or, with uv
uv run --with "maturin>=1.13,<2.0" maturin develop --release
```

### Release

The release process is automated via the `publish.yml` workflow triggered by version tags.

1. **Preconditions**:
   - PyPI project exists.
   - [Trusted Publisher](https://docs.pypi.org/trusted-publishers/) is configured for `publish.yml`.
   - GitHub environment `pypi` is configured (if repository is protected).

2. **Operator Flow**:
   Ensure the environment is ready and push a new tag:
   ```bash
   python -m pip install --upgrade pip
   pip install "maturin>=1.13,<2.0"
   git tag vX.Y.Z
   git push origin vX.Y.Z
   ```

The workflow builds cross-platform wheels and handles the PyPI upload automatically.

### Testing
Run the same local checks that CI runs:
```bash
source .venv/bin/activate

# Rust static validation
cargo fmt --all -- --check
cargo clippy --all-targets -- -D warnings
cargo test

# Python quality and release contract checks
basedpyright --project pyrightconfig.json
ruff check python tests scripts benchmarks
python scripts/check_release_contract.py metadata
python scripts/check_release_contract.py workflow --file .github/workflows/publish.yml

# Build the extension and run the Python suite
maturin develop --release
pytest tests/ -v --strict-markers -m "not stress"

# Optional longer determinism lane, matching the CI stress job
pytest tests/test_sort_false_determinism.py -v --strict-markers -m stress
```

With uv, run Python tools through the synced project environment:
```bash
uv sync --extra bench --extra dev
uv run --with "maturin>=1.13,<2.0" maturin develop --release
uv run basedpyright --project pyrightconfig.json
uv run ruff check python tests scripts benchmarks
uv run pytest tests/ -v --strict-markers -m "not stress"
```

### Benchmarking

#### Setup
To run benchmarks with library comparisons (Polars, etc.), install the optional benchmark dependencies:

```bash
# Install benchmark and development dependencies
source .venv/bin/activate
pip install -e ".[bench,dev]"

# Or with uv
uv sync --extra bench --extra dev

# Build the Rust extension in release mode
maturin develop --release
# Or with uv
uv run --with "maturin>=1.13,<2.0" maturin develop --release
```

This installs:
- `polars`: For performance comparison benchmarks
- `pyarrow`: For fast Polars -> Pandas conversion during correctness checks
- `matplotlib`: For generating plots and visualizations (optional)
- `pytest` and `pytest-benchmark`: For test framework and benchmarking

#### Running Benchmarks
With uv, replace `python ...` with `uv run python ...` in the commands below.

```bash
# Run default benchmarks (cardinality=all, diagnostic=none)
source .venv/bin/activate
python benchmarks/benchmark.py

# Run full suite (core + diagnostics)
python benchmarks/benchmark.py --cardinality all --diagnostic threshold --sort-mode unsorted

# Run only standard cardinality benchmarks
python benchmarks/benchmark.py --cardinality standard

# Run only high cardinality benchmarks
python benchmarks/benchmark.py --cardinality high

# Run only selected aggregation functions
python benchmarks/benchmark.py --agg std --agg var
python benchmarks/benchmark.py --agg median
python benchmarks/benchmark.py --agg prod
python benchmarks/benchmark.py --agg min --agg max --cardinality high --sort-mode sorted

# Add threshold-neighborhood diagnostics (opt-in)
python benchmarks/benchmark.py --diagnostic threshold --sort-mode unsorted

# Run only sorted or sort=False benchmarks
python benchmarks/benchmark.py --sort-mode sorted
python benchmarks/benchmark.py --sort-mode unsorted

# Combine options
python benchmarks/benchmark.py --cardinality high --sort-mode sorted

# Save per-aggregation benchmark reports
python benchmarks/benchmark.py --output benchmarks/reports

# Generate benchmark reports for all supported aggregations
python benchmarks/generate_docs.py

# Save internal single-key std/var profile evidence to JSON
python benchmarks/benchmark.py --agg std --agg var --profile-json profile.json

# Adjust sample count (applies to both cold and warm; default: 5)
python benchmarks/benchmark.py --samples 20
```

Note: `--agg` is repeatable and filters the benchmark to only the selected aggregation functions.
If omitted, benchmark reports default to `sum`. Single-key evidence sections are included in
Markdown reports only when selected `std`/`var` aggregations are emitted.

Note: `median` is fully supported by `--agg` selection for benchmark runs and correctness checks.
The dedicated `--profile-json` diagnostics remain focused on the single-key `std`/`var` evidence lane.

Note: `--profile-json` is an internal benchmark diagnostics output. By default it includes
single-key `std`/`var` evidence cases, plus phase breakdowns when a
Rust-only Booster profile hook is available. Cases that fall back to pandas or require Python
sorting remain in the JSON with `breakdown: null`.

Note: `--cardinality` is for workload classes (`standard`, `high`, `all`), while
`--diagnostic` is for internal boundary checks (`none`, `threshold`).

Note: `--diagnostic threshold` is only valid with `--sort-mode unsorted` because it targets
`sort=False` multi-key boundary behavior around `n_groups * n_keys ~= 200k`.

Breaking change (no compatibility mode): `--cardinality default` and
`--cardinality threshold` were removed in this version.

## Architecture Overview

`pandas-booster` uses a hybrid Rust/Python architecture:

- **PyO3**: Provides the bridge between Python and Rust.
- **Rayon**: Implements a work-stealing parallel scheduler for multi-core processing.
- **Radix Partitioning**: Multi-key groupby uses a 4-phase radix partitioning algorithm (histogram → prefix sum → scatter → aggregate) that eliminates merge overhead.
- **FixedKey Optimization**: For 1-10 key groupby operations (the supported maximum), uses compile-time fixed-size arrays (`FixedKey<const N>`) instead of dynamic vectors. This enables aggressive compiler optimizations (loop unrolling, SIMD) and avoids per-group heap allocation for key storage.
- **AHash**: Used for high-speed hashing of groupby keys.
- **SmallVec**: Used only as a generic fallback key representation (e.g., if the max-key constraint is raised in the future). The current implementation inlines up to 10 key values before spilling to the heap.
- **Zero-Copy**: NumPy arrays are accessed directly as Rust slices without copying data, minimizing memory overhead and latency.

The Python side provides a `BoosterAccessor` that handles validation and falls back to Pandas when the data doesn't meet the requirements for acceleration.

## License

This project is licensed under the MIT License.

