Metadata-Version: 2.4
Name: dqdetect
Version: 0.1.0
Summary: Command-line data quality profiler for CSV and Excel files
Author: Shahed
License-Expression: MIT
Project-URL: Homepage, https://github.com/Shahedr/data-quality-detective
Project-URL: Repository, https://github.com/Shahedr/data-quality-detective
Project-URL: Issues, https://github.com/Shahedr/data-quality-detective/issues
Keywords: data-quality,data-profiling,csv,excel,analytics,pandas
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas>=2.2
Requires-Dist: openpyxl>=3.1
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == "dev"
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: twine>=6.0; extra == "dev"
Dynamic: license-file

# Data Quality Detective

[![tests](https://github.com/Shahedr/data-quality-detective/actions/workflows/tests.yml/badge.svg)](https://github.com/Shahedr/data-quality-detective/actions/workflows/tests.yml) ![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue) ![MIT License](https://img.shields.io/badge/license-MIT-green)

A small command-line tool for quickly profiling CSV and Excel files before analysis.

I built this around a problem I keep running into in analytics work: a dataset can look usable at first, but the real issues only show up after checking missing values, duplicate rows, mixed numeric/text fields, date parsing, and unusual numeric values.

**dqdetect** puts those checks into one repeatable command and writes a report you can review or share.

## Quick start

~~~bash
git clone https://github.com/Shahedr/data-quality-detective.git
cd data-quality-detective

python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows
# .venv/Scripts/activate

pip install -e .
dqdetect examples/messy_orders.csv
~~~

By default the command writes both Markdown and HTML reports to the **reports/** folder.

## What it checks

- dataset shape and column types
- duplicate rows
- missing values by column
- constant columns
- mixed numeric/text values
- invalid values in date-like columns
- IQR-based outlier counts for numeric columns

The tool does **not** automatically delete or "fix" anything. The report is meant to help an analyst decide what deserves review.

## Example

~~~text
Data Quality Detective
File: examples/messy_orders.csv

Rows: 12
Columns: 8
Duplicate rows: 1

Reports written to:
  reports/messy_orders_report.md
  reports/messy_orders_report.html
~~~

See [examples/messy_orders_report.md](examples/messy_orders_report.md) for a sample report.

## CLI

~~~bash
dqdetect path/to/file.csv
dqdetect workbook.xlsx --sheet Orders\ndqdetect workbook.xlsx --sheet 0
dqdetect data.csv --format html
dqdetect data.csv --output audit_reports
~~~

Supported formats:

- CSV
- Excel (.xlsx)

## Why these checks?

This first version focuses on issues that can change an analysis if they are handled carelessly.

For example, a freight-cost field containing both dollar values and text such as "included elsewhere" should not simply be converted to zero. Likewise, an extreme numeric value may be a data-entry problem or a legitimate observation. The tool flags these cases instead of making the decision for you.

## Project structure

~~~text
data-quality-detective/
├── dqdetect/
│   ├── __init__.py
│   ├── checks.py
│   ├── cli.py
│   ├── io.py
│   ├── profiler.py
│   └── report.py
├── examples/
│   ├── messy_orders.csv
│   └── messy_orders_report.md
├── tests/
├── .github/workflows/tests.yml
├── CONTRIBUTING.md
├── LICENSE
└── pyproject.toml
~~~

## Development

~~~bash
pip install -e ".[dev]"
pytest
~~~

## Roadmap

Things I would like to add next:

- user-configurable thresholds
- JSON output for automation
- schema rules for expected columns and types
- PostgreSQL table profiling
- comparison between two versions of a dataset
- richer HTML charts
- optional GitHub Action for automated data-quality checks

I am keeping the first release intentionally small so the checks are understandable and easy to extend.

## License

MIT
