Metadata-Version: 2.4
Name: preventleak
Version: 0.1.0
Summary: Measure and prevent data leakage in machine learning evaluation
Author-email: Mohamed Aly Bouke <bouke@ieee.org>, Azizol Abdullah <azizol@upm.edu.my>, Nor Izura Udzir <izura@upm.edu.my>, Normalia Samian <normalia@upm.edu.my>, Mohamed Othman <mothman@upm.edu.my>
Maintainer-email: Mohamed Aly Bouke <bouke@ieee.org>
License-Expression: MIT
Project-URL: Homepage, https://pypi.org/project/preventleak/
Project-URL: Documentation, https://pypi.org/project/preventleak/
Keywords: data leakage,model evaluation,cross-validation,reproducibility,machine learning
Classifier: Programming Language :: Python :: 3
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.23
Requires-Dist: pandas>=1.5
Requires-Dist: scikit-learn>=1.1
Requires-Dist: scipy>=1.9
Requires-Dist: matplotlib>=3.5
Dynamic: license-file

# PreventLeak

**Measure and prevent data leakage in machine learning evaluation.**

Most tools tell you *whether* leakage is present. PreventLeak also tells you **how
much of your reported score is leakage and what the honest number is**, with a
calibrated confidence interval, and gives you leak-safe splits, folds, and cleaned
data to remove it. It works on any tabular feature matrix and any scikit-learn
compatible model, and it covers the *data-resident* leakage (duplicate, group,
temporal) that code-static analyzers cannot see.

- **Measure.** For every leakage channel, the optimism gap, the corrected (leak-free)
  score, a 95% confidence interval, and a significance test.
- **Prevent.** Leak-safe train/test splitting, cross-validation, duplicate removal,
  and a pipeline that fits every step on the training rows only.
- **Model-flexible.** Pass one model, several, or none (a documented default panel).
- **Any modality.** Tabular, and anything represented as features (text, image,
  time series).

---

## Table of contents

- [Installation](#installation)
- [Quickstart](#quickstart)
- [Concepts](#concepts)
  - [The optimism gap](#the-optimism-gap)
  - [Leakage channels](#leakage-channels)
  - [The confidence interval](#the-confidence-interval)
  - [An honest caveat: cross-validation pessimism](#an-honest-caveat-cross-validation-pessimism)
- [API reference](#api-reference)
  - [`PreventLeak`](#preventleakmodelnone-k_folds5-repeats3-seed00-modelsnone)
  - [`AuditReport`](#auditreport)
  - [`estimate_gap`](#estimate_gap)
  - [`LeakSpec`](#leakspec)
  - [`safe_split`](#safe_split)
  - [`safe_cv`](#safe_cv)
  - [`clean`](#clean)
  - [`LeakSafePipeline`](#leaksafepipeline)
  - [`auto_groups` and `detect_cross_split_duplicates`](#auto_groups-and-detect_cross_split_duplicates)
- [Command-line interface](#command-line-interface)
- [Reproducibility](#reproducibility)
- [Limitations](#limitations)
- [Citation](#citation)
- [License](#license)

---

## Installation

```bash
pip install preventleak
```

Requires Python >= 3.9 and `numpy`, `pandas`, `scikit-learn`, `scipy`, `matplotlib`
(installed automatically).

---

## Quickstart

### Measure the leakage optimism gap

```python
from preventleak import PreventLeak
from sklearn.ensemble import RandomForestClassifier

report = PreventLeak(RandomForestClassifier()).audit(
    X, y,
    channels=["scaling", "feature_selection", "mean_encoding"],
    cat_cols=[3],          # high-cardinality categorical -> target-encoding leakage
    group=patient_id,      # optional entity key -> group leakage
    time=timestamp,        # optional time order -> temporal leakage
)

print(report.summary())    # per-channel reported, corrected, gap, 95% CI, p
report.plot(save="audit.png")
records = report.to_dict() # machine-readable list of dicts
```

Example `summary()` output:

```
PreventLeak audit
channel                reported  Delta_hat         gap 95% CI corrected        p
mean_encoding[3]         0.7966     0.0089 [-0.0251, 0.0430]    0.7877    0.012
feature_selection        0.7744     0.0021 [-0.0365, 0.0408]    0.7723    0.695
scaling                  0.7943     0.0000 [-0.0345, 0.0346]    0.7943    0.746
-> largest optimism from 'mean_encoding[3]': reported 0.7966, honest 0.7877 ...
```

### Prevent it

```python
from preventleak import safe_split, safe_cv, clean, LeakSafePipeline
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import RandomForestClassifier

train_idx, test_idx = safe_split(X, y, group=patient_id, time=timestamp)

for tr, te in safe_cv(X, y, n_splits=5, group=patient_id):
    ...

X_clean, y_clean, kept = clean(X, y)   # drop cross-split duplicate rows

model = LeakSafePipeline(
    steps=[("scale", StandardScaler())],
    estimator=RandomForestClassifier(),
).fit(X[train_idx], y[train_idx])      # every step fit on train only
```

---

## Concepts

### The optimism gap

Let `S_leaky` be the score a practitioner reports under a protocol that contains a
leak, and `S_clean` the score of the same model under a leak-free protocol that
performs every data-dependent step inside the training portion of each fold. The
**leakage-induced optimism gap** is

```
Delta = S_leaky - S_clean
```

and the **corrected score** is `S_clean = S_leaky - Delta`. PreventLeak estimates
`Delta` for your specific dataset and pipeline, not its average over a benchmark, by
running your model under both protocols on the *same* cross-validation folds and
taking the paired difference.

### Leakage channels

A channel defines how the leaky and clean arms differ. Six are supported.

| channel | leaky arm | clean arm | discovered |
|---|---|---|---|
| `scaling` | standardizer fit on all rows | fit inside the fold | from pipeline |
| `feature_selection` | features ranked on all rows | ranked inside the fold | from pipeline |
| `mean_encoding` | category means on all rows | means inside the fold | from pipeline |
| `duplicate` | duplicate rows straddle the split | duplicates kept together | **automatic** |
| `group` | same entity on both sides | entity kept on one side | from `group=` key |
| `temporal` | train on past and future | train only on the past | from `time=` order |

Duplicate contamination is discovered without supervision; group and temporal
channels are activated when you pass a key or a time order; preprocessing channels
are activated for the steps you list.

### The confidence interval

Cross-validation folds share training rows, so the naive variance of the fold scores
is too small. PreventLeak uses the **Nadeau-Bengio corrected resampled-t variance**,
which inflates the naive variance by the test-to-train ratio `rho = 1/(k-1)` for
k-fold, so the interval reflects the true uncertainty of a cross-validation estimate.
Significance of the gap is a Wilcoxon signed-rank test on the paired fold scores.

### An honest caveat: cross-validation pessimism

The corrected score is a leak-free cross-validation estimate, so it carries the usual
cross-validation training-size pessimism: it estimates the performance of a model
trained on `(k-1)/k` of the data, which is slightly below a model trained on all of
it. This effect is a property of cross-validation, not of leakage; it shrinks as the
fold count rises, and PreventLeak does **not** count it as leakage.

---

## API reference

### `PreventLeak(model=None, k_folds=5, repeats=3, seed0=0, models=None)`

Audit a dataset for leakage-induced optimism.

- **model** - a single scikit-learn compatible estimator. Leave `None` to use a
  documented default panel (RandomForest(200), ExtraTrees(200), HistGradientBoosting).
- **models** - a dict `{name: estimator}` or a list of estimators to audit several at
  once; the report then carries a model column and `plot()` facets per model.
- **k_folds**, **repeats** - the repeated stratified k-fold design (default 5 x 3).
- **seed0** - base random seed.

**`.audit(X, y, channels=None, cat_cols=None, feature_k=10, group=None, time=None, detect_duplicates=True) -> AuditReport`**

- **X, y** - array-like features and binary target.
- **channels** - any of `"scaling"`, `"feature_selection"`, `"mean_encoding"`.
- **cat_cols** - column indices for `mean_encoding` (one channel per column).
- **feature_k** - number of features kept by `feature_selection`.
- **group** - per-row entity key; activates the `group` channel.
- **time** - per-row order or timestamp; activates the `temporal` channel.
- **detect_duplicates** - auto-discover duplicate contamination (default `True`).

### `AuditReport`

Returned by `audit`. Fields and methods:

- **`.results`** - list of `ChannelResult` (`channel`, `detected`, `gap`, `note`, `model`).
- **`.summary() -> str`** - a formatted per-channel table.
- **`.to_dict() -> list[dict]`** - one record per channel with `reported`, `corrected`,
  `gap`, `ci_low`, `ci_high`, `p_value`, `model`.
- **`.worst() -> ChannelResult`** - the channel with the largest gap.
- **`.plot(ax=None, alpha=0.05, title=..., save=None)`** - a forest plot of the
  per-channel gaps with 95% intervals; filled marks are significant. With several
  models it becomes small multiples.

Each `gap` is a `GapEstimate` with `reported`, `corrected`, `delta_hat`, `ci_low`,
`ci_high`, `wilcoxon_p`, `n_pairs`.

### `estimate_gap`

```python
estimate_gap(X, y, model, leak, groups=None, k_folds=5, repeats=5,
             seed0=0, rng_seed=0, metric="auc") -> GapEstimate
```

Low-level estimator for a single `LeakSpec`. `metric` is one of `"auc"`, `"acc"`,
`"f1"`, `"prauc"`, `"mcc"`.

### `LeakSpec`

```python
LeakSpec(kind, k=10, cat_col=-1, time=None)
```

Specifies one channel. `kind` is one of the six channel names; `k` is the feature
count for `feature_selection`; `cat_col` is the column for `mean_encoding`; `time`
is the per-row order for `temporal`.

### `safe_split`

```python
safe_split(X, y=None, test_size=0.25, group=None, time=None, dedup=True,
           stratify=True, random_state=0, decimals=6) -> (train_idx, test_idx)
```

A leak-free train/test split. A `time` order gives a forward split (train earlier,
test later); otherwise a `group` key or detected duplicates give a group-aware split
that never places rows of one group on both sides; otherwise a stratified split.

### `safe_cv`

```python
safe_cv(X, y=None, n_splits=5, group=None, time=None, dedup=True,
        random_state=0, decimals=6)  # yields (train_idx, test_idx)
```

Leak-free fold iterator: forward-chaining folds for `time`, group-disjoint folds for
a `group` key or detected duplicates, otherwise stratified folds.

### `clean`

```python
clean(X, y=None, dedup=True, decimals=6)
# -> (X_clean, y_clean, keep_mask)  or  (X_clean, keep_mask) when y is None
```

Drop exact or near-duplicate rows, keeping the first occurrence of each.

### `LeakSafePipeline`

```python
LeakSafePipeline(steps=None, estimator=None, dedup=True, decimals=6)
```

A thin wrapper that deduplicates the training rows and fits every transformer and the
estimator on the training data only. `steps` is a list of `(name, transformer)`;
use with `safe_cv` / `safe_split` for a fully leak-free evaluation. Provides `fit`,
`predict`, `predict_proba`.

### `auto_groups` and `detect_cross_split_duplicates`

```python
auto_groups(X, decimals=6) -> (group_ids, n_duplicate_rows)
detect_cross_split_duplicates(X_train, X_val, decimals=6) -> boolean_mask
```

`auto_groups` assigns a shared id to exact/near-duplicate rows; the second flags
validation rows whose match appears in the training portion.

---

## Command-line interface

Works on any CSV. Categorical columns are encoded automatically; the
highest-cardinality one is audited for target-encoding leakage.

```bash
# audit, save a plot and a JSON report
preventleak audit data.csv --target label --plot audit.png --json audit.json

# one or more models (shortcuts rf/hgb/et/logreg, or any dotted import path)
preventleak audit data.csv --target label --model rf logreg sklearn.ensemble.GradientBoostingClassifier

# with an entity key and a time column
preventleak audit data.csv --target label --group patient_id --time visit_date

# remove duplicate rows, or write a leak-safe split
preventleak clean data.csv --target label -o clean.csv
preventleak split data.csv --target label --time ts --out-prefix split
```

`audit` options: `--model` (one or more), `--folds`, `--repeats`, `--feature-k`,
`--group`, `--time`, `--plot`, `--json`.

---

## Reproducibility

The estimator is deterministic given the seeds (`seed0`, `rng_seed`). The default
model panel is fixed (RandomForest(200), ExtraTrees(200), HistGradientBoosting) so a
no-model audit is reproducible. Duplicate discovery rounds features to `decimals`
places before hashing.

---

## Limitations

- The corrected score inherits cross-validation training-size pessimism (see the
  caveat above); it is reported separately and is not counted as leakage.
- Group and temporal channels require you to supply an entity key or a time order;
  inferring them from features alone is not always possible.
- Preprocessing channels are audited for the steps you declare.
- Targets are treated as binary; multi-class support is on the roadmap.

---

## Citation

If you use PreventLeak, please cite:

> M. A. Bouke, A. Abdullah, N. I. Udzir, N. Samian, M. Othman.
> *Quantifying the Optimism Gap from Data Leakage in Machine Learning Evaluation.*

---

## License

MIT.
