Metadata-Version: 2.4
Name: tsfresh-rs
Version: 0.1.1
Classifier: Development Status :: 4 - Beta
Classifier: Programming Language :: Rust
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Developers
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Dist: numpy>=1.22
Requires-Dist: pandas>=1.3
Requires-Dist: scipy>=1.7
Requires-Dist: scikit-learn>=1.0 ; extra == 'sklearn'
Requires-Dist: pytest ; extra == 'test'
Requires-Dist: tsfresh>=0.21 ; extra == 'test'
Requires-Dist: scipy ; extra == 'test'
Requires-Dist: statsmodels ; extra == 'test'
Requires-Dist: mpmath ; extra == 'test'
Requires-Dist: scikit-learn ; extra == 'test'
Provides-Extra: sklearn
Provides-Extra: test
License-File: LICENSE
Summary: Fast, drop-in time series feature extraction for Python, powered by Rust
Keywords: time series,feature extraction,tsfresh,rust
Author: Fatin Ishraq
Requires-Python: >=3.10
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Changelog, https://github.com/Fatin-Ishraq/tsfresh-rs/releases
Project-URL: Homepage, https://github.com/Fatin-Ishraq/tsfresh-rs
Project-URL: Issues, https://github.com/Fatin-Ishraq/tsfresh-rs/issues
Project-URL: Repository, https://github.com/Fatin-Ishraq/tsfresh-rs

<p align="center">
  <img src="https://raw.githubusercontent.com/Fatin-Ishraq/tsfresh-rs/main/assets/logo.svg" alt="" width="104" height="104">
</p>

<h1 align="center">tsfresh-rs</h1>

<p align="center">
  <strong>Fast, drop-in time series feature extraction for Python, powered by Rust.</strong>
</p>

<p align="center">
  <a href="https://github.com/Fatin-Ishraq/tsfresh-rs/actions/workflows/ci.yml"><img src="https://github.com/Fatin-Ishraq/tsfresh-rs/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
  <a href="https://pypi.org/project/tsfresh-rs/"><img src="https://img.shields.io/pypi/v/tsfresh-rs.svg" alt="PyPI"></a>
  <a href="https://pypi.org/project/tsfresh-rs/"><img src="https://img.shields.io/badge/python-3.10%20--%203.14-blue.svg" alt="Python 3.10 to 3.14"></a>
  <a href="https://github.com/Fatin-Ishraq/tsfresh-rs/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-MIT-blue.svg" alt="License: MIT"></a>
</p>

## What it is

A machine that runs for a week produces a million sensor readings and no
columns. Most models want columns: one row per machine, one column per
property of its signal. Turning the first into the second is **time series
feature extraction**, and [`tsfresh`](https://github.com/blue-yonder/tsfresh)
is the standard Python library for it - it computes 783 such properties per
signal, from the mean to the Dickey-Fuller test statistic, then runs
hypothesis tests to tell you which ones actually relate to your target.

It is also slow enough that people give up on it. A comprehensive extraction
costs roughly **290 ms per series**, so a hundred thousand series is most of a
day.

`tsfresh-rs` is that library reimplemented in Rust. Same functions, same
arguments, same 783 columns under the same names in the same order - about
**100x faster**.

## What it does

Raw readings in, one row per series out:

```python
import numpy as np
import pandas as pd
import tsfresh_rs as tsfresh

rng = np.random.default_rng(0)
rows = []
for machine_id in range(20):
    signal = rng.standard_normal(100)
    if machine_id % 2:                       # half the machines drift
        signal += 0.03 * np.arange(100)
    rows.append(pd.DataFrame(
        {"id": machine_id, "time": np.arange(100), "vibration": signal}
    ))
df = pd.concat(rows, ignore_index=True)
```

```
   id  time  vibration
0   0     0   0.125730
1   0     1  -0.132105
2   0     2   0.640423
...
[2000 rows x 3 columns]
```

One call turns those 2,000 readings into a feature matrix:

```python
X = tsfresh.extract_features(df, column_id="id", column_sort="time")
```

```
X.shape
(20, 783)

   vibration__mean  vibration__standard_deviation  vibration__linear_trend__attr_"slope"
0           0.0811                         0.9621                                 0.0015
1           1.4344                         1.2601                                 0.0285
2          -0.1380                         1.1156                                -0.0030
3           1.4460                         1.2784                                 0.0306
4           0.0120                         1.0836                                -0.0040
5           1.4832                         1.1311                                 0.0241
```

783 columns is deliberately more than you need. If you have labels, a second
call keeps only the columns that are statistically related to them:

```python
from tsfresh_rs.utilities.dataframe_functions import impute

y = pd.Series([mid % 2 for mid in range(20)])   # which machines drifted
impute(X)                                        # selection rejects NaN
selected = tsfresh.select_features(X, y)
```

```
selected.shape
(20, 124)
```

That matrix goes straight into scikit-learn, or anything else that takes a
DataFrame.

## Install

```bash
pip install tsfresh-rs
```

Wheels for Linux, macOS and Windows. **Python 3.10 through 3.14** from a single
`abi3` wheel per platform, and nothing to compile.

## Already using tsfresh?

Change the import. That is the whole migration:

```diff
- import tsfresh
+ import tsfresh_rs as tsfresh
```

Same functions, same arguments, same 783 feature columns, same column names in
the same order - verified by 478 tests that run both libraries on the same
input and compare every value, including which inputs each one *refuses*.

If you cannot change the import - a pipeline that pulls in `tsfresh` by name
deep inside a dependency - alias the package instead:

```python
import tsfresh_rs
tsfresh_rs.install()   # before anything imports `tsfresh`

import tsfresh         # this is now tsfresh_rs
```

It refuses rather than half-patching the module graph if the real `tsfresh` has
already been imported.

## How much faster

<p align="center">
  <picture>
    <source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/Fatin-Ishraq/tsfresh-rs/main/assets/benchmark-dark.svg">
    <img src="https://raw.githubusercontent.com/Fatin-Ishraq/tsfresh-rs/main/assets/benchmark-light.svg" width="920" alt="End-to-end extraction speedup of tsfresh-rs over tsfresh: 93x to 141x across six workloads.">
  </picture>
</p>

| workload | `tsfresh` | `tsfresh-rs` | speedup |
|---|---:|---:|---:|
| 10 series × 500 pts | 2.94 s | 0.0278 s | **106x** |
| 20 series × 500 pts | 5.72 s | 0.0464 s | **123x** |
| 50 series × 500 pts | 14.18 s | 0.1007 s | **141x** |
| 100 series × 500 pts | 27.23 s | 0.1951 s | **140x** |

Two things are slow in the reference, and fixing only the obvious one gets you
a fraction of the win. The calculators themselves are slow - rewriting those in
Rust is worth about 17x. But most of a comprehensive extraction is not the
calculators at all: it is a Python call and a pandas `groupby`/`apply` per
series *per feature*. So this package does not call calculators one at a time.
It flattens every series into one contiguous buffer, compiles the whole feature
plan once, and makes a single call into Rust with the GIL released. Owning that
loop is what turns 17x into 120x.

Full tables - longer series, `n_jobs`, per calculator, and the preset that
exists to avoid the cost - are in **[docs/BENCHMARKS.md](docs/BENCHMARKS.md)**.

## What it adds: `feature_timings`

`tsfresh` ships `EfficientFCParameters`, a preset that drops
`approximate_entropy` and `sample_entropy` because they are expensive. That is a
guess - a judgement the maintainers made once, on their data. On a 500-point
series those two really are 62% of the run. On a 50-point series they are
nearly free. On a 20,000-point series something else dominates entirely.

`feature_timings` measures it on *your* data:

```python
import tsfresh_rs as tsfresh

timings = tsfresh.feature_timings(df, column_id="id", column_sort="time")
print(timings.head(6)[["feature", "columns", "seconds", "pct_of_total"]])
#                 feature  columns   seconds  pct_of_total
#     approximate_entropy        5  0.139937     60.287962
# augmented_dickey_fuller        3  0.032445     13.977916
#          sample_entropy        1  0.018846      8.119415
#        change_quantiles       60  0.007768      3.346718
#        number_cwt_peaks        2  0.007093      3.055826
#          ar_coefficient       11  0.006635      2.858552
```

The shares are exhaustive — all 74 features, summing to 100% — so nothing is
hiding outside the report.

And `drop_slowest` turns that report into a feature set:

```python
from tsfresh_rs.profiling import drop_slowest

cheap = drop_slowest(tsfresh.ComprehensiveFCParameters(), timings, budget_pct=90)
X = tsfresh.extract_features(df, default_fc_parameters=cheap,
                             column_id="id", column_sort="time")
```

`tsfresh` has no equivalent. The question "which of these 783 columns am I
paying for?" currently has no answer other than deleting features and
re-timing by hand.

## Is it actually the same?

That is the only question that matters for a drop-in, so it is the one the test
suite is built around.

- **478 tests**, almost all differential against `tsfresh` on the same input,
  comparing every value - and every refusal, so an input the reference rejects
  is rejected here too.
- **250 randomised fuzz cases** over length, scale, offset, tie density and
  degeneracy. This found four real defects during development, including a
  wrong Mexican-hat amplitude that scaled all 60 `cwt_coefficients` columns by
  1.3025 - a uniform error that looks entirely plausible in isolation.
- **API coverage is asserted mechanically**: every calculator in the reference
  must exist here with the same attributes, because those attributes are what
  define the presets.
- **Reference quirks are reproduced deliberately.** `fft_aggregated`'s kurtosis
  has a term the standardised fourth moment does not want;
  `ComprehensiveFCParameters` builds `mean_n_absolute_max` from a dict literal
  with three copies of one key, so two of its three intended features have
  never existed. Both are matched bug-for-bug, because a drop-in that quietly
  fixes a feature changes numbers people have already trained on.

A handful of values legitimately differ, and a few are *more* accurate here
than in the reference. Every one is documented and pinned by a test in
**[docs/COMPATIBILITY.md](docs/COMPATIBILITY.md)**.

## Limitations

Stated plainly, because a drop-in that hides its gaps is worse than one that
does not have them.

- **`matrix_profile` is unavailable**, exactly as in the reference. The
  `matrixprofile` package is unmaintained and does not build on current Python,
  so `tsfresh` ships with this feature disabled and excluded from
  `ComprehensiveFCParameters`. This mirrors that, including the `ImportError`.
- **`select_features` is not accelerated.** It is linear in the number of
  features with one cheap hypothesis test each — not the bottleneck this
  package exists to fix — so it runs in Python. `n_jobs` and `chunksize` are
  accepted and ignored there.
- **`chunksize`, `distributor` and the `profile*` arguments** describe
  tsfresh's multiprocessing pipeline, which this implementation does not have.
  They are accepted for signature compatibility and ignored, with a warning.
  The `utilities.distribution` classes exist so that code importing them keeps
  working, but `extract_features` does not route through them.
- **`tsfresh.__version__` reports this package's version after `install()`**,
  not a tsfresh one, so code gating on `tsfresh.__version__ >= "0.20"` will
  take the wrong branch. Claiming to be 0.21.2 would fix that one check and
  lie to every other; the API level being emulated is published separately as
  `tsfresh_rs.__tsfresh_version__`.
- **`profile=True` no longer answers "which calculator is slow".** A cProfile
  trace of an extraction here shows one opaque call into Rust. `feature_timings`
  is the replacement, and it measures what the profiler used to.
- **`augmented_dickey_fuller` is only 4x faster.** Its cost is a lag-selection
  search over many OLS fits, and the reference already spends that time inside
  compiled LAPACK.

## Development

```bash
pip install maturin pytest numpy pandas scipy tsfresh mpmath statsmodels
maturin build --release --out dist && pip install --no-index --find-links dist tsfresh-rs
pytest tests/ -q
python bench/bench.py
```

To reproduce the older-reference environment the version table describes — the
one where three feature columns legitimately differ — and check that the suite
skips exactly those:

```bash
uv venv --python 3.10 /tmp/old && uv pip install --python /tmp/old/bin/python     pytest numpy pandas scipy tsfresh mpmath statsmodels
uv pip install --python /tmp/old/bin/python --no-deps --find-links dist tsfresh-rs
/tmp/old/bin/python -m pytest tests/ -q -rs
```

That resolves to numpy 2.2, pandas 2.3, SciPy 1.15 and PyWavelets 1.8, and
gives `411 passed, 3 skipped`.

## Licence and credit

MIT, the same licence as `tsfresh`.

This is a reimplementation of [`tsfresh`](https://github.com/blue-yonder/tsfresh)
by Maximilian Christ and Blue Yonder GmbH, whose API, feature definitions and
output semantics it deliberately reproduces. If you use this in research, cite
their paper:

> Christ, M., Braun, N., Neuffer, J. and Kempa-Liehr A.W. (2018).
> *Time Series FeatuRe Extraction on basis of Scalable Hypothesis tests
> (tsfresh — A Python package).* Neurocomputing 307 (2018) 72-77.

