Metadata-Version: 2.4
Name: reproguard
Version: 0.2.0
Summary: Pre-production risk scanner for data science notebooks and ML repos.
Project-URL: Homepage, https://github.com/vipulgote1999/ReproGuard
Project-URL: Repository, https://github.com/vipulgote1999/ReproGuard
Project-URL: Changelog, https://github.com/vipulgote1999/ReproGuard/blob/main/CHANGELOG.md
Project-URL: Documentation, https://github.com/vipulgote1999/ReproGuard/blob/main/docs/GUIDES.md
Author: ReproGuard Contributors
License-Expression: MIT
License-File: LICENSE
Keywords: data-science,jupyter,leakage,mlops,privacy,reproducibility
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.10
Requires-Dist: jinja2>=3.1
Requires-Dist: nbclient>=0.10
Requires-Dist: nbformat>=5.10
Requires-Dist: packaging>=24.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: rich>=13.7
Requires-Dist: tomli>=2.0
Requires-Dist: typer>=0.12
Provides-Extra: dev
Requires-Dist: pytest-cov>=5.0; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Provides-Extra: privacy
Requires-Dist: presidio-analyzer>=2.2; extra == 'privacy'
Description-Content-Type: text/markdown

# ReproGuard

[![Python](https://img.shields.io/badge/python-3.10%2B-blue?logo=python&logoColor=white)](https://python.org)
[![License](https://img.shields.io/badge/license-MIT-green)](https://opensource.org/licenses/MIT)
[![Ruff](https://img.shields.io/badge/linted%20with-ruff-purple?logo=ruff)](https://docs.astral.sh/ruff)
[![Status](https://img.shields.io/badge/status-alpha-yellow?logo=databricks)](https://github.com/vipulgote1999/ReproGuard)

**Pre-production risk scanner for data science notebooks and ML repositories.**

ReproGuard analyzes Jupyter notebooks and Python scripts for reproducibility risks, data leakage, privacy violations, missing dependencies, and handoff readiness — before you share, review, or promote work toward production.

> **Positioning:** SonarQube-style review for data science work. Not a replacement for MLflow, DVC, Databricks, DataHub, or data observability platforms.

---

## Why This Exists

Data science projects often work on the author's machine but fail during review or production handoff because of:

- 📍 **Local data paths** — `C:\Users\...` or `/Users/...` that don't exist on another machine
- 📦 **Missing dependencies** — No `requirements.txt` or unpinned packages causing environment drift
- 🔄 **Out-of-order execution** — Notebook cells run in a non-linear order, hiding stateful assumptions
- 🔓 **Hidden PII & secrets** — Email addresses, API keys, or credentials buried in code or output
- 📊 **Data leakage** — Preprocessing before train/test split, test data in fit calls, target-like features
- 📝 **Missing handoff docs** — No clear statement of objective, data source, assumptions, or instructions

ReproGuard is **local-first** — scans happen on your machine without uploading data or notebooks to any third-party service.

---

## Features

### Detects 29 risk patterns across 6 categories

| Risk Category | What It Catches | Severity Range |
|---|---|---|
| **Reproducibility** | Missing dependency files, unpinned packages, no random seeds, out-of-order execution, stale outputs | LOW → HIGH |
| **Data Leakage** | Preprocessing before split, target-like feature columns, test data in fit calls, suspiciously high metrics, tabular data dumps, output tracebacks | LOW → CRITICAL |
| **Privacy & Security** | Email addresses, phone numbers, credit card numbers, AWS keys, hardcoded secrets, private key blocks, high-entropy credentials, base64 images, JSON blobs | LOW → CRITICAL |
| **Data Dependency** | Local machine paths, missing referenced data files, hardcoded paths in output | MEDIUM → HIGH |
| **Handoff Readiness** | Missing objective, data source, assumptions, or metric documentation | LOW |
| **Execution** | Notebook execution failures, kernel/dependency setup errors | HIGH → CRITICAL |

### Output Formats

- **Terminal** — Color-coded summary with severity breakdown
- **JSON** — Structured data for programmatic consumption
- **HTML** — Styled standalone report with issue grouping
- **SARIF 2.1.0** — Compatible with GitHub Code Scanning and VS Code

### Scoring

Penalty-based scoring from 0–100. Status: `ready_with_caution` (≥75), `needs_review` (50–74), or `not_ready` (<50 or any CRITICAL issue). All penalties and weights are configurable.

---

## Installation

```bash
pip install reproguard
```

### From source

```bash
git clone https://github.com/vipulgote1999/ReproGuard.git
cd ReproGuard
pip install -e .
```

### Development install

```bash
pip install -e .[dev]
```

### Privacy extra (Presidio support — future)

```bash
pip install reproguard[privacy]
```

---

## Quick Start

Scan a single notebook:

```bash
reproguard scan examples/risky_customer_churn.ipynb
```

Scan an entire project directory:

```bash
reproguard scan .
```

Scan with HTML report and fail CI on low score:

```bash
reproguard scan . --format html --fail-under 50
```

Enable clean notebook execution:

```bash
reproguard scan notebook.ipynb --execute
```

---

## Usage Examples

### Basic scan

```bash
reproguard scan examples/risky_customer_churn.ipynb
```

Output:

```
ReproGuard score: 27/100 (not_ready)
Files scanned: 1 | Issues: 8
Critical: 1 | High: 3 | Medium: 2 | Low: 2

 Severity   Code     Issue                          Location
 ────────   ────     ─────                          ────────
 CRITICAL   LEAK001  Possible preprocessing before…  risky_customer_churn.ipynb:cell 6
 HIGH       DATA001  Local machine path detected     risky_customer_churn.ipynb:line 2
 HIGH       LEAK002  Future/target-like column nam…  risky_customer_churn.ipynb:cell 4
 HIGH       LEAK007  Exception traceback found in…   risky_customer_churn.ipynb:cell 9
 MEDIUM     PII001   Email address detected          risky_customer_churn.ipynb:cell 8
 MEDIUM     REP001   Non-deterministic code witho…   risky_customer_churn.ipynb:cell 6
 LOW        NB001    Notebook cells were executed…   risky_customer_churn.ipynb
 LOW        NB002    Notebook output exists witho…   risky_customer_churn.ipynb:cell 9
```

### Generate reports

```bash
# All report formats
reproguard scan . --format all

# JSON only
reproguard scan . --format json

# SARIF for GitHub Code Scanning
reproguard scan . --format sarif

# Custom output directory
reproguard scan . --output-dir scan-reports
```

### CI integration

```bash
# Fail the build if the score is too low
reproguard scan . --fail-under 50
echo $?  # Exit code 1 when score < 50 or any CRITICAL issue
```

### Regression detection (baseline diff)

Gate CI on **new** issues while existing debt is paid down gradually:

```bash
# First run: save the baseline
reproguard scan . --format json --output-dir .reproguard

# Later runs: block on new issues only
reproguard scan . --baseline .reproguard/reproguard-report.json --fail-new 0
reproguard scan . --baseline .reproguard/reproguard-report.json --fail-new-critical
```

New and resolved issues are printed in the terminal summary and recorded in the JSON report metadata. Exit codes: `1` when the gate trips, `2` for invalid baselines.

### Scan with privacy disabled

```bash
reproguard scan . --no-privacy
```

### Custom execution timeout

```bash
reproguard scan notebook.ipynb --execute --execution-timeout 300
```

---

## Understanding Reports

### Score interpretation

| Score | Status | Action Required |
|-------|--------|----------------|
| ≥ 75 | `ready_with_caution` | Review minor issues before production |
| 50–74 | `needs_review` | Address significant issues |
| < 50 | `not_ready` | Blocking issues — must fix |
| Any CRITICAL | `not_ready` | Immediate attention required |
| 0 files scanned | `no_files` | No supported files found — check the scan path (exit code 2) |

### Report files

Reports are written to `.reproguard/` by default:

```
.reproguard/
├── reproguard-report.json      # Structured data
├── reproguard-report.html      # Styled HTML report
└── reproguard-report.sarif     # GitHub Code Scanning compatible
```

---

## Configuration

Create a `.reproguard.yml` in your project root:

```yaml
# .reproguard.yml
exclude_paths:
  - "archive/**"
  - "tests/**"
exclude_dirs:
  - "scratch"
fail_under: 50
checks:
  disabled:
    - "LEAK005"     # Disable large tabular output check
    - "PII004"      # Disable base64 image check
```

Configuration is discovered by walking up from the scan path (like git). See [docs/CONFIGURATION.md](docs/CONFIGURATION.md) for the full reference.

---

## Pre-commit Hook

```yaml
# .pre-commit-config.yaml
repos:
  - repo: https://github.com/vipulgote1999/ReproGuard
    rev: v0.2.0
    hooks:
      - id: reproguard-scan
        args: ["--fail-under", "75"]
```

The hook scans the entire repository on each commit and blocks the commit when the score falls below the threshold.

---

## CI/CD Integration

### GitHub Actions (with SARIF upload)

```yaml
name: ReproGuard
on: [push, pull_request]
jobs:
  scan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.11"
      - run: pip install reproguard
      - run: reproguard scan . --format sarif --fail-under 50
      - uses: github/codeql-action/upload-sarif@v3
        with:
          sarif_file: .reproguard/reproguard-report.sarif
```

### GitLab CI

```yaml
reproguard:
  stage: test
  script:
    - pip install reproguard
    - reproguard scan . --format html --fail-under 50
  artifacts:
    paths:
      - .reproguard/
```

---

## Project Structure

```
ReproGuard/
├── src/reproguard/
│   ├── cli.py              # Typer CLI entry point
│   ├── scanner.py          # Scan orchestrator
│   ├── models.py           # Core data models & scoring
│   ├── config.py           # .reproguard.yml loader
│   ├── python_analysis.py  # Python source analysis
│   ├── leakage.py          # ML data leakage heuristics
│   ├── privacy.py          # PII / secret scanning
│   ├── notebook.py         # Notebook parser
│   ├── dependency.py       # Dependency file analysis
│   ├── execution.py        # Clean notebook execution
│   ├── report.py           # JSON/HTML report generators
│   ├── sarif.py            # SARIF 2.1.0 report generator
│   ├── plugin.py           # Check registry & filtering
│   └── utils.py            # Shared helpers
├── examples/
│   ├── risky_customer_churn.ipynb  # Notebook with intentional issues
│   ├── clean_analysis.py           # Clean script example
│   └── requirements.txt            # Example dependency file
├── docs/
│   ├── ARCHITECTURE.md     # Internal design & module docs
│   ├── CHECKS.md           # Complete issue code reference
│   ├── CONFIGURATION.md    # Configuration file reference
│   ├── GUIDES.md           # Usage guides & integrations
│   └── ROADMAP.md          # Future plans
├── pyproject.toml          # Build & tool config
└── README.md               # This file
```

---

## Check Reference

| Code | Check | Severity | Category |
|------|-------|----------|----------|
| PY001 | Python syntax error | CRITICAL | Reproducibility |
| DATA001 | Local machine path | HIGH | Data Dependency |
| DATA002 | Referenced data file not found | MEDIUM | Data Dependency |
| REP001 | Non-deterministic code without seed | MEDIUM | Reproducibility |
| LEAK001 | Preprocessing before train/test split | CRITICAL | Data Leakage |
| LEAK002 | Future/target-like column name | HIGH | Data Leakage |
| LEAK003 | Test data in fitting call | CRITICAL | Data Leakage |
| LEAK004 | Suspiciously high metric | MEDIUM | Data Leakage |
| LEAK005 | Large tabular output | LOW | Data Leakage |
| LEAK006 | Hardcoded path in output | MEDIUM | Data Dependency |
| LEAK007 | Exception traceback in output | HIGH | Reproducibility |
| PII001 | Email address | HIGH | Privacy |
| PII002 | Phone number | MEDIUM | Privacy |
| PII003 | Credit card number | HIGH | Privacy |
| PII004 | Base64 image in output | LOW | Privacy |
| PII005 | Large JSON/blob in output | LOW | Privacy |
| SEC001 | AWS access key | CRITICAL | Privacy |
| SEC002 | Hardcoded secret assignment | CRITICAL | Privacy |
| SEC003 | Private key block | CRITICAL | Privacy |
| SEC004 | High-entropy string | MEDIUM | Privacy |
| NB001 | Out-of-order execution | MEDIUM | Reproducibility |
| NB002 | Output without execution count | LOW | Reproducibility |
| NB003 | Non-default kernel requirement | MEDIUM | Reproducibility |
| DEP001 | No dependency file found | HIGH | Reproducibility |
| DEP002 | Unpinned dependency | LOW | Reproducibility |
| DEP003 | Imported package missing from deps | MEDIUM | Reproducibility |
| HAND001 | Missing handoff documentation | LOW | Handoff |
| EXEC001 | Notebook execution error | CRITICAL | Execution |
| EXEC002 | Execution setup failure | HIGH | Execution |

See [docs/CHECKS.md](docs/CHECKS.md) for full details on every check.

---

## Design Principles

1. **Local-first** — Scans run entirely on your machine. No data or notebooks leave your environment.
2. **Explainable rules** — Every issue has a code, severity, evidence, confidence score, and suggested fix. No black boxes.
3. **Low friction** — CLI-first design with a single `reproguard scan <target>` command. Pre-commit hook, CI integration, and GitHub Action out of the box.
4. **Conservative scoring** — The score is transparent (penalty-based, weighted by severity and category). You can customize all penalties and weights.
5. **Narrow wedge** — Focused on catching pre-production risks before work enters heavier MLOps pipelines.

---

## Limitations

ReproGuard v0.1 (alpha) uses heuristics. It flags likely risks but cannot prove every issue is real. Treat it as a **review assistant**, not a final governance decision.

- Leakage detection is heuristic — expect false positives (use `checks.ignore` to silence known-safe locations)
- Dependency parsing is intentionally lightweight (regex-based for requirements/Pipfile, YAML for conda env files, TOML for pyproject/lock files — no full resolver)
- Privacy scanning uses regex rules by default (Presidio support planned)
- Clean notebook execution depends on local kernel and dependency availability
- Data file existence checks are limited to paths referenced from notebooks/scripts

---

## Development

```bash
# Install dev dependencies
pip install -e .[dev]

# Lint
ruff check .

# Test
pytest

# Build release artifacts and validate metadata
python -m build
python -m twine check dist/*
```

Releases follow the procedure in [docs/RELEASING.md](docs/RELEASING.md) — see
the checklist there before tagging. Changes are tracked in
[CHANGELOG.md](CHANGELOG.md).

---

## Roadmap

- **v0.2** ✅ — correctness hardening (magic handling, conda/lock dependency
  parsing), path-scoped ignore rules, baseline diffing, extended coverage
- **v0.3**: GitHub Action, custom rule engine, parallel notebook execution,
  kernel-aware `--execute`
- **v1.0+**: ML pipeline scanning, differential scans, team dashboard, API

See [docs/ROADMAP.md](docs/ROADMAP.md) for the full roadmap.

---

## License

MIT
