Metadata-Version: 2.4
Name: chemdatacheck
Version: 0.1.0
Summary: Quality checks for molecular datasets: integrity, duplicates, leakage, ML readiness, and chemical space
Author: Alessio Prunotto
License-Expression: MIT
Project-URL: Homepage, https://github.com/AlessioPrunotto/chemdatacheck
Project-URL: Documentation, https://github.com/AlessioPrunotto/chemdatacheck/blob/main/docs/tutorial.md
Project-URL: Repository, https://github.com/AlessioPrunotto/chemdatacheck
Project-URL: Issues, https://github.com/AlessioPrunotto/chemdatacheck/issues
Project-URL: Changelog, https://github.com/AlessioPrunotto/chemdatacheck/blob/main/CHANGELOG.md
Keywords: cheminformatics,molecular-machine-learning,dataset-quality,data-leakage,rdkit
Classifier: Development Status :: 3 - Alpha
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Chemistry
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: rdkit>=2022.9.5
Requires-Dist: pandas>=1.5
Requires-Dist: numpy>=1.23
Requires-Dist: scipy>=1.9
Requires-Dist: pyarrow>=10.0.1
Requires-Dist: openpyxl>=3.1
Provides-Extra: pretty
Requires-Dist: rich>=13; extra == "pretty"
Provides-Extra: dev
Requires-Dist: pytest>=7.4; extra == "dev"
Requires-Dist: pytest-cov>=5; extra == "dev"
Requires-Dist: ruff>=0.6; extra == "dev"
Provides-Extra: notebooks
Requires-Dist: jupyter>=1; extra == "notebooks"
Requires-Dist: nbformat>=5.9; extra == "notebooks"
Requires-Dist: nbclient>=0.8; extra == "notebooks"
Requires-Dist: ipykernel>=6.25; extra == "notebooks"
Dynamic: license-file

# ChemDataCheck

**Quality checks for molecular datasets — like pytest for the data behind
chemistry and molecular-machine-learning projects.**

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![CI](https://github.com/AlessioPrunotto/chemdatacheck/actions/workflows/ci.yml/badge.svg)](https://github.com/AlessioPrunotto/chemdatacheck/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/chemdatacheck.svg)](https://pypi.org/project/chemdatacheck/)
<!-- Uncomment after the first GitHub release + Zenodo hookup:
[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.XXXXXXX.svg)](https://doi.org/10.5281/zenodo.XXXXXXX)
-->

![ChemDataCheck turns molecular datasets into prioritized, evidence-backed findings across chemical integrity, duplicates, split leakage, ML readiness, and chemical-space coverage.](docs/assets/chemdatacheck-graphical-abstract.svg)

One command to sanity-check a chemistry dataset before you train, publish, or trust it:

```bash
python -m pip install chemdatacheck
chemdatacheck dataset.csv
```

For datasets with labels and predefined train/test splits:

```bash
chemdatacheck dataset.csv --label-col activity --split-col split \
  --format html --output chemdatacheck-report.html
```

It checks for:
 - valid molecules
 - duplicates
 - stereochemical collisions
 - salt/tautomer
 - ambiguity
 - split leakage
 - target shift
 - chemical-space bias
 - etc.

Every warning comes with row IDs, evidence, and an actionable recommendation,
plus a 0–100 dataset quality score. See an
[example HTML report](examples/curated_report.html), the
[scientific validation](docs/validation.md), and the
[interpretation limits](#what-chemdatacheck-does-not-do). Questions and field
reports are welcome in the
[field-report form](https://github.com/AlessioPrunotto/chemdatacheck/issues/new?template=field-report.yml).

## Install

ChemDataCheck requires Python 3.10 or newer and is tested with Python 3.10–3.12.
Install it from PyPI in an isolated environment. On macOS or Linux:

```bash
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install chemdatacheck
```

On Windows PowerShell, create and activate the environment with:

```powershell
py -3.11 -m venv .venv
.venv\Scripts\Activate.ps1
python -m pip install chemdatacheck
```

To add colored terminal output, install `chemdatacheck[pretty]` instead.
Required dependencies—including RDKit, pandas, SciPy, PyArrow, and
openpyxl—are declared in `pyproject.toml` and installed automatically by pip.

If a compatible RDKit wheel is not available for your Python version or
platform, use a Conda environment instead:

```bash
conda create -n chemdatacheck -c conda-forge python=3.11 rdkit pip
conda activate chemdatacheck
python -m pip install chemdatacheck
```

Contributors who need an editable installation, tests, coverage reporting, or
the notebooks should follow [`CONTRIBUTING.md`](CONTRIBUTING.md).

## Inputs

CSV/TSV, SDF/SD, Parquet, Excel, JSON-lines. The SMILES column is autodetected
(`smiles`, `canonical_smiles`, ...); alternatively, you can pass `--smiles-col`. Splits (training set / test set)
can be passed via a column (`--split-col`) or by passing two files (`train.csv test.csv`).

## Exit codes (pytest-like)

- `0` clean
- `1` warnings
- `2` errors

## What it checks

- **Chemical integrity:** invalid SMILES, valence errors, aromaticity/kekulization,
  impossible charges, disconnected components (salts/mixtures), isotopes,
  radicals, unspecified stereocenters, tautomer ambiguity.
- **Dataset duplicates:** exact, canonical, stereochemical collisions, salt/solvate
  duplicates, tautomer duplicates, near-duplicates (Morgan fingerprints, similarity
  measured with Tanimoto distance).
- **Split leakage:** 2D-identity leakage, analog leakage (Tc ≥ 0.6), near-duplicates
  across splits, scaffold overlap, suspiciously-easy-split heuristic.
- **ML readiness:** duplicated/conflicting measurements, target distribution shift,
  split-predicts-label leakage suspects, label outliers.
- **Chemical space:** rare elements, rare/promiscuous functional groups, unusual
  ring systems, representation bias, applicability-domain gaps.

Every finding carries row IDs, evidence examples, and a recommendation, e.g.
`train_row` ↔ `test_row` pairs with Tanimoto and shared scaffold for leakage.

To run leakage checks, ChemDataCheck must know which rows are training data and
which are held out for testing or validation. Common split values such as
`train`, `test`, and `valid` are recognized automatically. If your dataset uses
different names, map them explicitly; for example:

```bash
chemdatacheck dataset.csv --split-col partition \
  --train-value development --test-value external
```

For numbered cross-validation folds, you only need to identify the held-out
fold. This treats fold `0` as test data and every other observed fold as
training data:

```bash
chemdatacheck dataset.csv --split-col fold --test-value 0
```

If ChemDataCheck cannot form both a non-empty training group and a non-empty test
group, it reports a warning and does not run the leakage checks. If only some
split values are mapped, it warns that the remaining rows were excluded from
those checks.

The quality score is a prioritization heuristic, not a validated scientific metric.
The quality score starts from 100, and points are subtracted for each bad quality finding.
Related findings are overlap-capped so the same underlying
invalid structure, duplicate, or leakage issue is not fully deducted multiple
times. JSON reports record tool/RDKit versions, settings, resolved columns, and
SHA-256 input hashes for reproducibility.
When a large-dataset check uses sampling or a bounded candidate search, JSON
reports expose it in `meta.approximations` and in the finding's `metadata`.

## What ChemDataCheck does not do

ChemDataCheck is an evidence-producing audit and triage tool, not an automatic
certification or data-cleaning system. It does not silently rewrite structures,
prove that a model will generalize, or replace assay review and domain expertise.
A flagged salt, tautomer, scaffold overlap, or repeated measurement may be
scientifically appropriate; the row-level evidence is more important than the
summary score.

## Learn

- **Tutorial:** [`docs/tutorial.md`](docs/tutorial.md) — from install to CI-gated
  audits in ~20 minutes.
- **Scientific validation:** [`docs/validation.md`](docs/validation.md) — public
  ML datasets plus fixed ChEMBL/PubChem samples, seeded defects, threshold
  sensitivity, false-positive guidance, and runtime through 40k+ molecules.
- **Examples:** [`examples/`](examples/) — three executed notebooks, all data
  generated inline (no downloads):
  - `01_quickstart.ipynb` — CLI + Python API on a deliberately dirty dataset.
  - `02_leakage_splits.ipynb` — identity vs analog vs scaffold leakage, and fixes.
  - `03_curation_case_study.ipynb` — full curation loop (39 → 80) with HTML report.
