Metadata-Version: 2.5
Name: veridelta
Version: 0.22.0
Summary: Compare two datasets on their primary keys under rules you declare, on a laptop, in CI, or inside a warehouse.
Project-URL: Homepage, https://veridelta.github.io/veridelta/
Project-URL: Repository, https://github.com/Veridelta/veridelta
Project-URL: Issues, https://github.com/Veridelta/veridelta/issues
Author-email: The Veridelta Contributors <veridelta.labs@gmail.com>
License: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: bigquery,data-comparison,data-diff,data-engineering,data-migration,databricks,duckdb,mcp,polars,snowflake
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Database
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Software Development :: Testing
Classifier: Typing :: Typed
Requires-Python: >=3.11
Requires-Dist: polars>=1.39.3
Requires-Dist: pydantic>=2.12.5
Requires-Dist: pyyaml>=6.0.3
Requires-Dist: typing-extensions>=4.14.1
Provides-Extra: all
Requires-Dist: connectorx>=0.4.5; extra == 'all'
Requires-Dist: databricks-sql-connector>=3.0.0; extra == 'all'
Requires-Dist: deltalake>=0.15.0; extra == 'all'
Requires-Dist: duckdb>=1.5.0; extra == 'all'
Requires-Dist: fastexcel>=0.11.0; extra == 'all'
Requires-Dist: google-cloud-bigquery>=3.11.0; extra == 'all'
Requires-Dist: mcp<3,>=2.3.0; extra == 'all'
Requires-Dist: pyarrow<24,>=14.0.1; (python_version >= '3.14') and extra == 'all'
Requires-Dist: pyarrow>=14.0.1; extra == 'all'
Requires-Dist: pyarrow>=14.0.1; (python_version < '3.14') and extra == 'all'
Requires-Dist: pyiceberg>=0.6.0; extra == 'all'
Requires-Dist: rapidfuzz>=3.0.0; extra == 'all'
Requires-Dist: snowflake-connector-python>=4.7.3; extra == 'all'
Provides-Extra: bigquery
Requires-Dist: google-cloud-bigquery>=3.11.0; extra == 'bigquery'
Requires-Dist: pyarrow>=14.0.1; extra == 'bigquery'
Provides-Extra: database
Requires-Dist: connectorx>=0.4.5; extra == 'database'
Requires-Dist: pyarrow>=14.0.1; extra == 'database'
Provides-Extra: databricks
Requires-Dist: databricks-sql-connector>=3.0.0; extra == 'databricks'
Requires-Dist: pyarrow>=14.0.1; extra == 'databricks'
Provides-Extra: delta
Requires-Dist: deltalake>=0.15.0; extra == 'delta'
Provides-Extra: duckdb
Requires-Dist: duckdb>=1.5.0; extra == 'duckdb'
Requires-Dist: pyarrow>=14.0.1; extra == 'duckdb'
Provides-Extra: excel
Requires-Dist: fastexcel>=0.11.0; extra == 'excel'
Provides-Extra: fuzzy
Requires-Dist: rapidfuzz>=3.0.0; extra == 'fuzzy'
Provides-Extra: iceberg
Requires-Dist: pyiceberg>=0.6.0; extra == 'iceberg'
Provides-Extra: mcp
Requires-Dist: mcp<3,>=2.3.0; extra == 'mcp'
Provides-Extra: snowflake
Requires-Dist: pyarrow<24,>=14.0.1; (python_version >= '3.14') and extra == 'snowflake'
Requires-Dist: pyarrow>=14.0.1; (python_version < '3.14') and extra == 'snowflake'
Requires-Dist: snowflake-connector-python>=4.7.3; extra == 'snowflake'
Description-Content-Type: text/markdown

# Veridelta

[![CI Pipeline](https://github.com/veridelta/veridelta/actions/workflows/ci.yml/badge.svg)](https://github.com/veridelta/veridelta/actions)
[![codecov](https://codecov.io/gh/veridelta/veridelta/graph/badge.svg)](https://codecov.io/gh/veridelta/veridelta)
[![PyPI version](https://badge.fury.io/py/veridelta.svg)](https://pypi.org/project/veridelta/)
[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)

Veridelta compares two datasets on their primary keys and reports every row that differs under the rules you declare. Nothing is forgiven unless a rule says so, and the exit code tells CI whether the datasets match. Use it to verify a migration, a pipeline change, or a model's new evaluation run, on a laptop, in CI, or inside a warehouse.

![A terminal prints a five-line veridelta.yaml and two three-row CSV files, validates the configuration, runs the comparison, shows one added, one removed, and one changed row, and prints the exit code for CI, 1, beside what 0, 1, and 3 mean.](https://veridelta.github.io/veridelta/assets/demo.gif)

It runs on [Polars](https://pola.rs/). Read the [documentation](https://veridelta.github.io/veridelta/).

## Features

- **Declared rules.** Tolerances, null sentinels, regular expressions, value maps, date parsing, casts, and fuzzy text matching apply in nine fixed stages. Nothing is forgiven unless a rule says so, and `strict_types` fails a column whose type drifts.
- **Comparison inside the warehouse.** Two tables in Snowflake, Databricks, or BigQuery are compared where they are stored, as are two Postgres or DuckDB tables that set `pushdown`. The rules compile to SQL, and only counts and keys come back. [Pushdown](https://veridelta.github.io/veridelta/pushdown/) lists the exceptions and the services the SQL has run in.
- **Many sources.** CSV, Parquet, JSON, NDJSON, Arrow, Avro, and Excel files, Delta Lake and Iceberg tables, DuckDB files and MotherDuck databases, and Postgres, MySQL, SQL Server, Oracle, SQLite, and other databases. Files and tables are scanned lazily where Polars can.
- **Built for CI.** Exit codes, a JSON summary, a standalone HTML report, a Markdown summary for pull requests, OpenTelemetry metrics, and files of the rows that differ. A GitHub Action posts the summary on each pull request, and a GitLab CI template on each merge request when it has a token.
- **Checks before a run.** `veridelta validate` reports what would stop a run without reading any rows, and a JSON Schema gives editors completion for configuration files.
- **For AI agents.** An agent runs the same checks and comparisons through the command line, or through `veridelta mcp`, a Model Context Protocol server that reads files only from the folders you name. [AI agents](https://veridelta.github.io/veridelta/agents/) gives the steps.

## Install

```bash
uv add veridelta                # or: pip install veridelta
uv add 'veridelta[snowflake]'   # extras: snowflake, databricks, bigquery, delta, iceberg, database, duckdb, excel, fuzzy, mcp, all
```

## Quick start

The smallest configuration names the two files and the keys that pair their rows. The suffix of each path says what format it is:

```yaml
# veridelta.yaml
primary_keys: [id]
source:
  path: legacy.csv
target:
  path: modern.csv
```

The recording above runs this file on two three-row files. `validate` checks the file without reading any rows, and `run` compares the two files. It exits 0 when they match, 1 when rows differ, and 3 when the run could not finish:

```bash
veridelta validate -c veridelta.yaml
veridelta run -c veridelta.yaml
```

Rules say what counts as a match, column by column. This file forgives one percent on a total and compares phone numbers on their digits alone:

```yaml
primary_keys: ["transaction_id"]
source:
  path: "legacy.parquet"
target:
  path: "modern.parquet"
rules:
  - column_names: ["grand_total"]
    relative_tolerance: 0.01
  - column_names: ["contact_number"]
    regex_replace: {"[^0-9]": ""}
```

In Python, `DiffEngine` compares two `LazyFrame`s with the same models:

```python
import polars as pl
from veridelta import DiffConfig, DiffEngine, DiffRule

result = DiffEngine(
    DiffConfig(
        primary_keys=["user_id"],
        rules=[DiffRule(pattern="^AMT_.*", absolute_tolerance=0.05)],
    ),
    pl.scan_parquet("legacy.parquet"),
    pl.scan_parquet("modern.parquet"),
).run()

if not result.summary.is_match:
    raise SystemExit(f"{result.summary.changed_count} rows differ")
```

## Documentation

- [Tutorials](https://veridelta.github.io/veridelta/examples/01_core_concepts/): five notebooks, from a first comparison in Python to a CI pipeline.
- [User guide](https://veridelta.github.io/veridelta/configuration/): configuration, sources, rules, pushdown, results, the command line, and AI agents.
- [CI integrations](https://veridelta.github.io/veridelta/ci/): the GitHub Action and the GitLab CI template.
- [API reference](https://veridelta.github.io/veridelta/api/): the public Python interface.
- [Roadmap](https://veridelta.github.io/veridelta/roadmap/): work that is not built yet.

## Accessibility

[ACCESSIBILITY.md](https://github.com/Veridelta/veridelta/blob/main/ACCESSIBILITY.md) states what Veridelta aims for, the barriers known today, and how to report one.

## Contributing

See [CONTRIBUTING.md](https://github.com/Veridelta/veridelta/blob/main/CONTRIBUTING.md) for the development setup and the checks a change must pass.

## License

Apache 2.0.
