Metadata-Version: 2.4
Name: tsxtract-rs
Version: 0.3.0
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Rust
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Typing :: Typed
Requires-Dist: numpy>=1.24
Requires-Dist: pandas>=1.5 ; extra == 'bench'
Requires-Dist: tsfresh>=0.20 ; extra == 'bench'
Requires-Dist: pycatch22>=0.4.5 ; extra == 'bench'
Requires-Dist: tsfel>=0.1.6 ; extra == 'bench'
Requires-Dist: mkdocs>=1.6 ; extra == 'docs'
Requires-Dist: mkdocs-material>=9.5 ; extra == 'docs'
Requires-Dist: pandas>=1.5 ; extra == 'pandas'
Requires-Dist: pytest>=7.0 ; extra == 'test'
Requires-Dist: scipy>=1.10 ; extra == 'test'
Requires-Dist: hypothesis>=6.100 ; extra == 'test'
Requires-Dist: pandas>=1.5 ; extra == 'test'
Provides-Extra: bench
Provides-Extra: docs
Provides-Extra: pandas
Provides-Extra: test
License-File: LICENSE
Summary: Fast time-series feature extraction, Rust core
Keywords: time-series,feature-extraction,rust,tsfresh,machine-learning
License-Expression: MIT
Requires-Python: >=3.10
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Changelog, https://github.com/Aamod007/Tsxtract/blob/main/CHANGELOG.md
Project-URL: Homepage, https://github.com/Aamod007/Tsxtract
Project-URL: Issues, https://github.com/Aamod007/Tsxtract/issues
Project-URL: Repository, https://github.com/Aamod007/Tsxtract

# tsxtractor

Batch time-series feature extraction for Python, with a Rust core.

33 curated statistical, temporal, and spectral features, computed across a whole
batch of series at once. The Rust core takes zero-copy views of your numpy
buffers, releases the GIL, and parallelises across the *series* dimension with
[rayon](https://github.com/rayon-rs/rayon) — so throughput scales with your
cores when you have many series to process.

```bash
pip install tsxtractor
```

Wheels ship for Linux (x86_64, aarch64), macOS (Intel, Apple Silicon), and
Windows (x86_64), for Python 3.10+. No Rust toolchain needed.

## Quickstart

```python
import numpy as np
import tsxtractor

# batch: one row per series
X = np.random.randn(100_000, 500)
feats = tsxtractor.extract_features(X)        # (100_000, 33) float64
names = tsxtractor.feature_names()            # stable column order

# same thing with labeled columns (needs pandas)
df = tsxtractor.extract_features_df(X)

# ragged series of different lengths
feats = tsxtractor.extract_features([arr1, arr2, arr3])

# rolling windows over one long series
feats = tsxtractor.sliding_features(x, window=256, stride=64)
```

Input must be float64 and C-contiguous — a wrong dtype raises `TypeError`
instead of being silently copied, so you decide where the conversion cost is
paid (`X.astype(np.float64)`).

## Where the speed comes from

Parallelism is across **series**, not across the feature computations within one
series. The practical consequences:

- Extracting from 100 000 series scales close to linearly with core count.
- Extracting from *one* series shows no speedup versus a good numpy
  implementation, and is not meant to. If your workload is one long series, use
  `sliding_features`, which parallelises over the windows.

The GIL is released for the whole compute region, so this parallelism is real
under threaded Python callers, not capped by the interpreter lock.

## Benchmark

```bash
pip install -e ".[bench]"
python benches/bench_libraries.py --n-series 1000 --n-steps 500
```

| library | features | total time | series/s | ms/feature | vs tsxtractor |
|---|---:|---:|---:|---:|---:|
| **tsxtractor 0.2.1** | 33 | **1.2 ms** | 800,256 | 0.0379 | baseline |
| `catch22` (pycatch22) | 22 | 1.02 s | 976 | 46.58 | 820x slower |
| `TSFEL` (all domains) | 156 | 7.15 s | 140 | 45.86 | 5,725x slower |
| `tsfresh` (EfficientFCParameters) | 777 | 17.68 s | 57 | 22.76 | 14,151x slower |

1000 series x 500 steps, 16 cores, Windows 11, Python 3.14. Each library gets
the best of as many runs as fit in a two-second budget. Feature counts differ,
so `ms/feature` is the column to read for like-for-like work; `series/s` is the
one that decides how long your batch job takes.

The benchmark compares batch throughput against `tsfresh`, `catch22`, and
`TSFEL` on the same input, timing each library end to end *including* the input
reshaping it requires (tsfresh needs a long DataFrame; catch22 and TSFEL need a
per-series Python loop). It is run on a GitHub Linux runner by the
[Benchmark workflow](.github/workflows/benchmark.yml) so the published numbers
are reproducible by a stranger rather than measured on one laptop; the latest
committed run is in [`benches/results/`](benches/results/).

Two things drive the gap, and only one of them is engineering:

- **Batch parallelism.** Every other library here is called once per series from
  Python, so the batch cost is a serial loop plus per-call overhead. tsxtractor
  takes the whole matrix across the FFI boundary once and spreads the series
  over cores.
- **A cheaper feature set.** All 33 features are O(n) or O(n log n) by
  construction. `catch22` includes costlier estimators, which is why it stays
  slower per feature even on a single series: on one 500-point series,
  `catch22_all` takes 945 us (43 us/feature) against 9.4 us (0.28 us/feature)
  for `extract_features`, both measured as best of 200 runs. That is a
  difference in what is being computed, not only in how fast it is computed —
  if you need those specific estimators, this table is not telling you to
  switch.

## When not to use this

- **You need exhaustive feature coverage.** `tsfresh` computes up to 1 558
  features and `TSFEL` around 390. tsxtractor computes 33, chosen to stay
  low-redundancy. If you want to throw everything at a feature selector, use
  `tsfresh`.
- **You need custom or parameterised features.** The feature set is
  intentionally closed; there is no plugin hook.
- **You are in R, Julia, or MATLAB.** Use `catch22`, which has bindings for all
  three. tsxtractor is Python-only.
- **Your workload is one short series at a time.** The parallelism has nothing
  to work with; numpy is fine.

## Comparison

| Library | Features | Core | Batch-parallel | Notes |
|---|---:|---|---|---|
| **tsxtractor** | 33 | Rust + PyO3 | yes (rayon, across series) | curated, low-redundancy set; numpy-only dependency |
| `tsfresh` | up to 1 558 | Python | no | most exhaustive; slowest by a wide margin at scale |
| `TSFEL` | ~390 | Python | no | broad domain coverage; high within-set redundancy |
| `catch22` | 22 | C | no | fixed, literature-selected set; multi-language bindings |
| `tsflex` | n/a | Python | no | windowing framework, not a feature bank — calls others |

## Features (33)

| Group | Features |
|---|---|
| Stats | mean, std, var, min, max, median, quantile_10/25/75/90, skewness, kurtosis, abs_energy, root_mean_square |
| Change | mean_abs_change, mean_change, cid_ce (z-normalized), mean_second_derivative_central |
| Counts | zero_crossings, mean_crossings, number_of_peaks (support 3), longest_strike_above/below_mean |
| Correlation | autocorr at lags 1, 2, 5, 10; linear trend slope and r² |
| Entropy | permutation_entropy (order 3, normalized to [0, 1]) |
| Spectral | dominant_frequency, spectral_centroid, spectral_entropy (positive bins, DC excluded, sample spacing 1) |

Conventions: population moments (`ddof=0`); numpy-default linear interpolation
for quantiles; skewness/kurtosis follow `scipy.stats` with `bias=True`
(kurtosis is Fisher/excess).

`feature_names()` order is a **stability guarantee**: column `i` means the same
feature for every release within a major version. Reordering or renaming is a
major-version change.

## NaN policy

Two separate categories, deliberately not conflated:

- **NaN is a value.** Any NaN anywhere in a series makes all 33 of that series'
  features NaN — no silent imputation. Features that are individually undefined
  for an otherwise-valid series (autocorrelation or spectral features of a
  constant series, change features of a length-1 series) are NaN on their own
  while the rest compute normally.
- **Structural problems raise.** No series at all, a zero-length series, a
  non-contiguous array, `window`/`stride` < 1, or `window` longer than the
  series raise `ValueError`. A wrong dtype or shape raises `TypeError`. No Rust
  panic crosses the boundary; this is enforced by property-based tests.

## Correctness

Every feature is checked against a numpy/scipy reference implementation across
normal, trending, periodic, constant, two-element, single-element, and
heavily-tied series. CI publishes a
[reference-validation report](docs/validation.md) with the max absolute error
per feature, plus `hypothesis` property tests asserting that no input produces a
panic, a wrong output shape, or a NaN-policy violation.

## Development

```bash
python -m venv .venv && . .venv/bin/activate   # .venv\Scripts\activate on Windows
pip install maturin
pip install -e ".[test]"
maturin develop --release
pytest tests/
cargo test --no-default-features    # pure-Rust unit tests
```

See [CONTRIBUTING.md](CONTRIBUTING.md) for how to add a feature and what
requires a version bump.

## License

MIT — see [LICENSE](LICENSE).

