Metadata-Version: 2.5
Name: veridelta
Version: 0.14.10
Summary: Compare two datasets under declared rules, locally or inside the warehouse.
Project-URL: Homepage, https://veridelta.github.io/veridelta/
Project-URL: Repository, https://github.com/Veridelta/veridelta
Project-URL: Issues, https://github.com/Veridelta/veridelta/issues
Author-email: The Veridelta Contributors <veridelta.labs@gmail.com>
License: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: data-engineering,diff,polars,testing
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Typing :: Typed
Requires-Python: >=3.11
Requires-Dist: polars>=1.39.3
Requires-Dist: pydantic>=2.12.5
Requires-Dist: pyyaml>=6.0.3
Provides-Extra: all
Requires-Dist: connectorx>=0.4.5; extra == 'all'
Requires-Dist: databricks-sql-connector>=3.0.0; extra == 'all'
Requires-Dist: deltalake>=0.15.0; extra == 'all'
Requires-Dist: duckdb>=1.5.0; extra == 'all'
Requires-Dist: fastexcel>=0.11.0; extra == 'all'
Requires-Dist: google-cloud-bigquery>=3.11.0; extra == 'all'
Requires-Dist: pyarrow<24,>=14.0.1; (python_version >= '3.14') and extra == 'all'
Requires-Dist: pyarrow>=10.0.1; extra == 'all'
Requires-Dist: pyarrow>=10.0.1; (python_version < '3.14') and extra == 'all'
Requires-Dist: pyarrow>=14.0.1; extra == 'all'
Requires-Dist: pyiceberg>=0.6.0; extra == 'all'
Requires-Dist: rapidfuzz>=3.0.0; extra == 'all'
Requires-Dist: snowflake-connector-python>=3.8.0; extra == 'all'
Provides-Extra: bigquery
Requires-Dist: google-cloud-bigquery>=3.11.0; extra == 'bigquery'
Requires-Dist: pyarrow>=14.0.1; extra == 'bigquery'
Provides-Extra: database
Requires-Dist: connectorx>=0.4.5; extra == 'database'
Requires-Dist: pyarrow>=14.0.1; extra == 'database'
Provides-Extra: databricks
Requires-Dist: databricks-sql-connector>=3.0.0; extra == 'databricks'
Requires-Dist: pyarrow>=10.0.1; extra == 'databricks'
Provides-Extra: delta
Requires-Dist: deltalake>=0.15.0; extra == 'delta'
Provides-Extra: duckdb
Requires-Dist: duckdb>=1.5.0; extra == 'duckdb'
Requires-Dist: pyarrow>=14.0.1; extra == 'duckdb'
Provides-Extra: excel
Requires-Dist: fastexcel>=0.11.0; extra == 'excel'
Provides-Extra: fuzzy
Requires-Dist: rapidfuzz>=3.0.0; extra == 'fuzzy'
Provides-Extra: iceberg
Requires-Dist: pyiceberg>=0.6.0; extra == 'iceberg'
Provides-Extra: snowflake
Requires-Dist: pyarrow<24,>=14.0.1; (python_version >= '3.14') and extra == 'snowflake'
Requires-Dist: pyarrow>=10.0.1; (python_version < '3.14') and extra == 'snowflake'
Requires-Dist: snowflake-connector-python>=3.8.0; extra == 'snowflake'
Description-Content-Type: text/markdown

# Veridelta

[![CI Pipeline](https://github.com/veridelta/veridelta/actions/workflows/ci.yml/badge.svg)](https://github.com/veridelta/veridelta/actions)
[![codecov](https://codecov.io/gh/veridelta/veridelta/graph/badge.svg)](https://codecov.io/gh/veridelta/veridelta)
[![PyPI version](https://badge.fury.io/py/veridelta.svg)](https://pypi.org/project/veridelta/)
[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)

Veridelta compares two datasets on their primary keys and reports every row that differs once the rules you declare are applied. Use it to verify a system migration, a model retrain, or a pipeline change.

It runs on [Polars](https://pola.rs/). Read the [documentation](https://veridelta.github.io/veridelta/).

## Features

- **Declared rules.** Tolerances, null sentinels, regular expressions, value maps, date parsing, casts, and fuzzy text matching apply in nine fixed stages. Nothing is forgiven unless a rule says so, and `strict_types` fails a column whose type drifts.
- **The same verdict in the warehouse.** Two tables in Snowflake, Databricks, BigQuery, Postgres, or DuckDB are compared where they are stored. The rules compile to SQL, and only counts and keys come back. [Pushdown](https://veridelta.github.io/veridelta/pushdown/) lists the exceptions.
- **Many sources.** CSV, Parquet, JSON, Arrow, Avro, and Excel files, Delta Lake and Iceberg tables, DuckDB files and MotherDuck databases, and Postgres, MySQL, SQL Server, Oracle, SQLite, and other databases. Files and tables are scanned lazily where Polars can.
- **Built for CI.** Exit codes, a JSON summary, a standalone HTML report, a Markdown summary for pull requests, OpenTelemetry metrics, and files of the rows that differ. A GitHub Action and a GitLab CI template post the summary on each pull request.
- **Checks before a run.** `veridelta validate` reports what would stop a run without reading any rows, and a JSON Schema gives editors completion for configuration files.

## Install

```bash
uv add veridelta                # or: pip install veridelta
uv add 'veridelta[snowflake]'   # extras: snowflake, databricks, bigquery, delta, iceberg, database, duckdb, excel, fuzzy, all
```

## Quick start

In Python, `DiffEngine` compares two `LazyFrame`s:

```python
import polars as pl
from veridelta import DiffConfig, DiffEngine, DiffRule

result = DiffEngine(
    DiffConfig(
        primary_keys=["user_id"],
        rules=[DiffRule(pattern="^AMT_.*", absolute_tolerance=0.05)],
    ),
    pl.scan_parquet("legacy.parquet"),
    pl.scan_parquet("modern.parquet"),
).run()

if not result.summary.is_match:
    raise SystemExit(f"{result.summary.changed_count} rows differ")
```

The same comparison as a YAML file, for the CLI and CI:

```yaml
# veridelta.yaml
primary_keys: ["transaction_id"]
source:
  path: "legacy.parquet"
  format: "parquet"
target:
  path: "modern.parquet"
  format: "parquet"
rules:
  - column_names: ["grand_total"]
    relative_tolerance: 0.01
  - column_names: ["contact_number"]
    regex_replace: {"[^0-9]": ""}
```

`validate` checks the file without reading any rows, and `run` compares the datasets:

```bash
veridelta validate -c veridelta.yaml
veridelta run -c veridelta.yaml
```

## Documentation

- [Tutorials](https://veridelta.github.io/veridelta/examples/01_core_concepts/): five notebooks, from a first comparison in Python to a CI pipeline.
- [User guide](https://veridelta.github.io/veridelta/configuration/): configuration, sources, rules, pushdown, results, and the command line.
- [CI integrations](https://veridelta.github.io/veridelta/ci/): the GitHub Action and the GitLab CI template.
- [API reference](https://veridelta.github.io/veridelta/api/): the public Python interface.
- [Roadmap](https://veridelta.github.io/veridelta/roadmap/): work that is not built yet.

## Accessibility

[ACCESSIBILITY.md](https://github.com/Veridelta/veridelta/blob/main/ACCESSIBILITY.md) states what Veridelta aims for, the barriers known today, and how to report one.

## Contributing

See [CONTRIBUTING.md](https://github.com/Veridelta/veridelta/blob/main/CONTRIBUTING.md) for the development setup and the checks a change must pass.

## License

Apache 2.0.
