Metadata-Version: 2.4
Name: data-quality-agent-dk
Version: 0.1.1
Summary: Agentic Data Quality Pipeline: profiles CSVs, proposes DQ rules, writes+validates SQL to check them, and reports confirmed violations.
Author: Dev Khanna
Project-URL: Homepage, https://github.com/dev-khanna/data_quality_agent_dk
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: langchain>=0.3
Requires-Dist: langgraph>=0.2
Requires-Dist: duckdb>=1.0
Requires-Dist: python-dotenv>=1.0
Requires-Dist: pydantic>=2.0
Requires-Dist: streamlit>=1.35
Requires-Dist: pandas>=2.0
Requires-Dist: typing_extensions>=4.6
Provides-Extra: google
Requires-Dist: langchain-google-genai>=2.0; extra == "google"
Provides-Extra: anthropic
Requires-Dist: langchain-anthropic>=0.2; extra == "anthropic"
Provides-Extra: openai
Requires-Dist: langchain-openai>=0.2; extra == "openai"

# data-quality-agent-dk

Agentic data-quality pipeline. Point it at a folder of CSVs; it profiles
every table, proposes DQ rules from real evidence, writes + validates the
SQL to check each one (read-only, sandboxed, retried on failure), and
hands back a JSONL report plus an optional Streamlit dashboard.

## Requirements

- Python 3.10+
- An API key for at least one LLM provider (Examples include Gemini, OpenRouter, OpenAI)

## Installation

Running just:

```bash
pip install data-quality-agent-dk
```

gets you the base package — langchain, langgraph, duckdb, streamlit, and
everything else the pipeline itself needs. What it does **not** include
is a library for talking to any specific LLM provider, since most people
only use one and there's no reason to install libraries for providers
you're not using.

So say you want to use Gemini: you add `[google]` to the install command,
which tells pip "install the base package, and also install what's
needed for Gemini support":

```bash
pip install data-quality-agent-dk[google]
```

That `[google]` is called an **extra** — an optional add-on dependency
set. Under the hood it just pulls in `langchain-google-genai` alongside
everything else. The same idea applies for the other providers:


Any model + provider that langchain's 'init_chat_model' supports will 
work (Eg: OpenRouter). Install the respectibe provider's own package yourself
and pass it as a parameter while calling 'run_pipeline' below. (see Step 2 below)


## Step 1 — Set up your environment

Create a `.env` file in your project's working directory (not inside this
package) with whichever credentials your chosen models need. At minimum,
the LLM provider key(s) for whatever you pass as `planning_model_provider`
/ `worker_model_provider`:

```
# pick whichever provider(s) you're using
GOOGLE_API_KEY=...
ANTHROPIC_API_KEY=...
OPENAI_API_KEY=...
OPENROUTER_API_KEY=...

# optional: enables per-table LangSmith tracing (profiling -> rule
# planning -> every SQL generate/execute retry, as one trace per table)
LANGSMITH_TRACING_V2=true
LANGSMITH_API_KEY=...
LANGSMITH_PROJECT=your-project-name
```

You only need the key(s) matching the provider(s) you actually pass in —
if both `planning_model_provider` and `worker_model_provider` are
`"google_genai"`, you only need `GOOGLE_API_KEY`.

## Step 2 — Run it

```python
from data_quality_agent_dk import run_pipeline

result = run_pipeline(
    data_dir="data",              # folder of .csv files, one table per file
    output_dir=".",               # where dq_report.jsonl / todo_list.md land

    # any model string + provider langchain's init_chat_model accepts
    planning_model="gemini-2.5-flash",
    planning_model_provider="google_genai",
    worker_model="gemini-2.5-flash-lite",
    worker_model_provider="google_genai",

    temperature=1.0,
    max_tokens=40000,

    sample_rows_limit=100,
    max_retries=3,
    max_violation_rows_shown=3,
    suspicious_violation_ratio=0.45,

    display=True,   # launch the Streamlit dashboard when the run finishes
)

print(result["report_path"], result["failed_tables"])
```

`data_dir` is the only required argument. Everything else has a default
(see the table below) — but `planning_model`/`planning_model_provider`
and `worker_model`/`worker_model_provider` default to Gemini models, so
if you're using a different provider, **you must set those four
explicitly**, or the pipeline will try to call Gemini regardless of which
API key you've set.

### What you get back

`run_pipeline` returns a dict:

| Key | What it is |
|---|---|
| `table_names` | Every table (CSV) that was processed |
| `failed_tables` | Tables that errored out and were skipped |
| `report_path` | Path to `dq_report.jsonl` — one line per confirmed violation |
| `todo_path` | Path to `todo_list.md` — every rule considered, with its final status |

If `display=True` (the default), a Streamlit dashboard also opens
automatically once the run finishes, showing the same report visually
with filters and sample violating rows per rule.

## Choosing your own values

Every field below has a working default — you only need to touch a
value if your data or budget genuinely calls for something different.

| Parameter | Ask yourself | Default |
|---|---|---|
| `planning_model` | This makes one exhaustive, column-by-column reasoning pass per table — is my data complex enough (many columns, many edge cases) to justify a stronger/pricier model here? | `gemini-2.5-flash` |
| `worker_model` | This is called once *per rule* — PK inference, every SQL write/repair, every report entry. It runs many times per table, so cost and latency matter more than raw reasoning depth here. Do I want the cheapest/fastest model that's still reliable? | `gemini-2.5-flash-lite` |
| `temperature` | Do I want the same table to produce roughly the same rules and SQL if I re-run it? Lower = more deterministic/reproducible. Higher = more variation between runs. | `1.0 (gemini supports this better)` |
| `max_tokens` | How wide are my tables, and how long do my rule descriptions/SQL tend to get? A table with 40+ columns or many cross-column rules needs more headroom than a 5-column table. | `40000` |
| `sample_rows_limit` | How big is my dataset, and how representative does a sample need to be to catch rare-but-real issues? A 5000-row table can be mostly sampled; a 5-million-row table needs a bigger absolute sample to surface low-frequency problems, even though it's a smaller fraction of the whole. | `100` |
| `max_retries` | How much do I value catching every possible rule vs. how much do I care about runtime/cost? Higher retries recover more rules from a bad first attempt, at the cost of more LLM calls per rule that's genuinely hard to express in SQL. | `3` |
| `max_violation_rows_shown` | Am I using this report for a human to skim, or feeding it into another system? Fewer sample rows keeps the report/dashboard readable; more gives more diagnostic context per issue. | `3` |
| `suspicious_violation_ratio` | How dirty do I expect this data to genuinely be? Raise it if you have reason to expect genuinely high violation rates. | `0.45` |
| `display` | Toggling to true allows you to view the output in a simple streamlit interface | `True` |

For the full technical reference — every config field, the graph
architecture, and troubleshooting — see [`docs/config.md`](docs/config.md).
