Metadata-Version: 2.4
Name: reproguard
Version: 0.3.1
Summary: Pre-production risk scanner for data science notebooks and ML repos.
Project-URL: Homepage, https://github.com/vipulgote1999/ReproGuard
Project-URL: Repository, https://github.com/vipulgote1999/ReproGuard
Project-URL: Documentation, https://github.com/vipulgote1999/ReproGuard/blob/main/docs/CHECKS.md
Project-URL: Changelog, https://github.com/vipulgote1999/ReproGuard/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/vipulgote1999/ReproGuard/issues
Author: Vipul Gote
License: REPROGUARD SOFTWARE LICENSE (RSL)
        Version 1.0
        Copyright (c) 2026 Vipul Gote. All rights reserved.
        
        This is a proprietary source-available license. It is NOT an open-source
        license. Read it carefully before using the Software.
        
        1. OWNERSHIP
           ReproGuard — including its source code, compiled artifacts, documentation,
           design, architecture, algorithms, heuristics, check definitions, and all
           related materials (the "Software") — is the exclusive intellectual property
           of Vipul Gote. All rights not expressly granted in this license are
           reserved to Vipul Gote.
        
        2. GRANT OF USE
           Vipul Gote grants you a free, non-exclusive, non-transferable, revocable,
           personal license to USE the Software, in unmodified form, solely for its
           intended purpose: pre-production risk scanning of data science notebooks,
           Python scripts, and machine-learning repositories.
        
        3. PERMITTED USE
           You may install and run the Software, and use it to scan projects you are
           authorized to scan, subject to the terms of this license. This grant covers
           personal, academic, and internal business use of the scanning tool itself.
        
        4. PROHIBITED ACTS
           You may NOT, without the prior written permission of Vipul Gote:
             a) copy, reproduce, duplicate, or clone the Software or any portion of
                it beyond what is technically necessary to install and run it;
             b) modify, adapt, translate, merge, or create derivative works from the
                Software;
             c) distribute, publish, sublicense, rent, lease, lend, or sell the
                Software or any copy of it, in whole or in part;
             d) use the Software, its source code, its ideas, or its check logic to
                build, develop, or power a competing product, tool, plugin, or
                service;
             e) remove, alter, or obscure any copyright, trademark, or proprietary
                notice contained in the Software;
             f) claim authorship, ownership, or inventorship of the Software or its
                underlying ideas;
             g) resell access to the Software or offer it as a hosted service
                (SaaS) without prior written permission.
        
        5. ATTRIBUTION
           Where the Software is displayed, referenced, or demonstrated (including
           in screenshots, blog posts, talks, or derived documentation), you must
           credit it as "ReproGuard by Vipul Gote".
        
        6. TERMINATION
           This license is effective until terminated. It terminates automatically
           and immediately upon any breach of its terms. Upon termination you must
           cease all use of the Software and destroy all copies in your possession.
        
        7. NO WARRANTY
           THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS
           OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF
           MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, AND NON-INFRINGEMENT.
           THE ENTIRE RISK OF USING THE SOFTWARE RESTS WITH YOU.
        
        8. LIMITATION OF LIABILITY
           IN NO EVENT SHALL VIPUL GOTE BE LIABLE FOR ANY CLAIM, DAMAGES, OR OTHER
           LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT, OR OTHERWISE, ARISING
           FROM, OUT OF, OR IN CONNECTION WITH THE SOFTWARE OR ITS USE.
        
        9. GOVERNING LAW
           This license shall be governed by and construed in accordance with the
           laws of the Republic of India. Any dispute arising under this license
           shall be subject to the exclusive jurisdiction of the courts of India.
        
        10. ENTIRE AGREEMENT
            This license constitutes the entire agreement between you and Vipul Gote
            regarding the Software and supersedes all prior agreements or
            understandings. Any permission granted by Vipul Gote must be in writing
            to be effective.
License-File: LICENSE
Keywords: data-science,jupyter,leakage,mlops,privacy,reproducibility
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: Other/Proprietary License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.10
Requires-Dist: jinja2>=3.1
Requires-Dist: nbclient>=0.10
Requires-Dist: nbformat>=5.10
Requires-Dist: packaging>=24.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: rich>=13.7
Requires-Dist: tomli>=2.0
Requires-Dist: typer>=0.12
Provides-Extra: dev
Requires-Dist: pytest-cov>=5.0; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Provides-Extra: privacy
Requires-Dist: presidio-analyzer>=2.2; extra == 'privacy'
Description-Content-Type: text/markdown

# ReproGuard

[![Python](https://img.shields.io/badge/python-3.10%2B-blue?logo=python&logoColor=white)](https://python.org)
[![License](https://img.shields.io/badge/license-proprietary-blue)](https://github.com/vipulgote1999/ReproGuard/blob/main/LICENSE)
[![Copyright](https://img.shields.io/badge/copyright-%C2%A9%202026%20Vipul%20Gote-blue)](https://github.com/vipulgote1999/ReproGuard/blob/main/LICENSE)
[![Ruff](https://img.shields.io/badge/linted%20with-ruff-purple?logo=ruff)](https://docs.astral.sh/ruff)
[![PyPI](https://img.shields.io/pypi/v/reproguard?color=blue&logo=pypi&logoColor=white)](https://pypi.org/project/reproguard)
[![Status](https://img.shields.io/badge/status-alpha-yellow?logo=databricks)](https://github.com/vipulgote1999/ReproGuard)
[![Tests](https://img.shields.io/badge/tests-351%20passing-green?logo=pytest)](https://github.com/vipulgote1999/ReproGuard/actions)

**Pre-production risk scanner for data science notebooks and ML repositories.**

ReproGuard analyzes Jupyter notebooks and Python scripts for reproducibility risks, data leakage, privacy violations, missing dependencies, and handoff readiness — before you share, review, or promote work toward production.

> **Positioning:** SonarQube-style review for data science work. Not a replacement for MLflow, DVC, Databricks, DataHub, or data observability platforms.

---

## What's New in v0.3

v0.3 doubles the check inventory (29 → **73**) and adds a dedicated **GenAI**
risk category plus a full supporting toolchain:

| Area | What's new |
|------|------------|
| **GenAI category (27 checks)** | LLM client config (temperature, max_tokens, model pinning, timeouts), prompt hygiene (injection bait, untrusted interpolation, duplicates), LangGraph/CrewAI agent structure, tool agency & HITL, system-prompt leakage, fine-tuning contamination — aligned with [OWASP LLM Top 10 2025](https://genai.owasp.org/) |
| **Repository hygiene (REPO001-007)** | `.env` committed without gitignore, missing `.env.example`, large committed model artifacts, tests without CI, missing READMEs, DVC outputs missing from disk |
| **Notebook hidden state (NB004-006)** | Use-before-assign across cells, undefined names, code-only notebooks with no narrative |
| **Unsafe Python (SEC005-011)** | `eval`/`exec`, unsafe `yaml.load`, shell subprocesses, SQL string interpolation, HTTP without timeouts, assert outside tests |
| **Custom rule engine** | Your own regex checks with full registry integration (`checks.disabled`/`checks.ignore` work on them) |
| **Scan profiles** | `--profile genai` — category weights tuned per project type |
| **Auto-fixes** | `--fix gitignore|clear-counts|seed|all` — safe, idempotent remediations |
| **Git-history secret scan** | gitleaks-lite over `git log -p` using the PII/secret pipeline |
| **LLM config manifests** | Validate `llm.yaml`/`llm.json` (model pinning, providers, prompt paths; optional online catalog check) |
| **Model cards** | Markdown handoff artifact per scan |
| **Prompt artifact scanning** | `prompts/*.md|txt|yaml|json` scanned as whole prompts |
| **Execution upgrades** | Parallel notebook execution + kernel-aware `--execute` |
| **GitHub Action** | Reusable composite action for CI |

Every feature is documented with runnable examples below.

---

## Why This Exists

Data science projects often work on the author's machine but fail during review or production handoff because of:

- 📍 **Local data paths** — `C:\Users\...` or `/Users/...` that don't exist on another machine
- 📦 **Missing dependencies** — No `requirements.txt` or unpinned packages causing environment drift
- 🔄 **Out-of-order execution** — Notebook cells run in a non-linear order, hiding stateful assumptions
- 🔓 **Hidden PII & secrets** — Email addresses, API keys, or credentials buried in code or output
- 📊 **Data leakage** — Preprocessing before train/test split, test data in fit calls, target-like features
- 📝 **Missing handoff docs** — No clear statement of objective, data source, assumptions, or instructions

ReproGuard is **local-first** — scans happen on your machine without uploading data or notebooks to any third-party service.

---

## Features

### Detects 73 risk patterns across 7 categories

| Risk Category | What It Catches | Severity Range |
|---|---|---|
| **Reproducibility** | Missing dependency files, unpinned packages, no random seeds, out-of-order execution, stale outputs | LOW → HIGH |
| **Data Leakage** | Preprocessing before split, target-like feature columns, test data in fit calls, suspiciously high metrics, tabular data dumps, output tracebacks | LOW → CRITICAL |
| **Privacy & Security** | Email addresses, phone numbers, credit card numbers, AWS keys, hardcoded secrets, private key blocks, high-entropy credentials, base64 images, JSON blobs | LOW → CRITICAL |
| **Data Dependency** | Local machine paths, missing referenced data files, hardcoded paths in output | MEDIUM → HIGH |
| **Handoff Readiness** | Missing objective, data source, assumptions, or metric documentation | LOW |
| **Execution** | Notebook execution failures, kernel/dependency setup errors | HIGH → CRITICAL |
| **GenAI** | LLM client config (temperature, max_tokens, model pinning, timeouts), prompt injection bait, interpolated untrusted content, LangGraph/CrewAI agent structure, tool agency, system-prompt leakage, fine-tuning contamination, llm.yaml manifests | LOW → CRITICAL |

### Output Formats

- **Terminal** — Color-coded summary with severity breakdown
- **JSON** — Structured data for programmatic consumption
- **HTML** — Styled standalone report with issue grouping
- **SARIF 2.1.0** — Compatible with GitHub Code Scanning and VS Code

### Scoring

Penalty-based scoring from 0–100. Status: `ready_with_caution` (≥75), `needs_review` (50–74), or `not_ready` (<50 or any CRITICAL issue). All penalties and weights are configurable.

---

## Installation

```bash
pip install reproguard
```

### From source

```bash
git clone https://github.com/vipulgote1999/ReproGuard.git
cd ReproGuard
pip install -e .
```

### Development install

```bash
pip install -e .[dev]
```

### Privacy extra (Presidio support — future)

```bash
pip install reproguard[privacy]
```

---

## Quick Start

Scan a single notebook:

```bash
reproguard scan examples/risky_customer_churn.ipynb
```

Scan an entire project directory:

```bash
reproguard scan .
```

Scan with HTML report and fail CI on low score:

```bash
reproguard scan . --format html --fail-under 50
```

Enable clean notebook execution:

```bash
reproguard scan notebook.ipynb --execute
```

---

## Usage Examples

### Basic scan

```bash
reproguard scan examples/risky_customer_churn.ipynb
```

Output:

```
ReproGuard score: 27/100 (not_ready)
Files scanned: 1 | Issues: 8
Critical: 1 | High: 3 | Medium: 2 | Low: 2

 Severity   Code     Issue                          Location
 ────────   ────     ─────                          ────────
 CRITICAL   LEAK001  Possible preprocessing before…  risky_customer_churn.ipynb:cell 6
 HIGH       DATA001  Local machine path detected     risky_customer_churn.ipynb:line 2
 HIGH       LEAK002  Future/target-like column nam…  risky_customer_churn.ipynb:cell 4
 HIGH       LEAK007  Exception traceback found in…   risky_customer_churn.ipynb:cell 9
 MEDIUM     PII001   Email address detected          risky_customer_churn.ipynb:cell 8
 MEDIUM     REP001   Non-deterministic code witho…   risky_customer_churn.ipynb:cell 6
 LOW        NB001    Notebook cells were executed…   risky_customer_churn.ipynb
 LOW        NB002    Notebook output exists witho…   risky_customer_churn.ipynb:cell 9
```

### Scan a GenAI project

ReproGuard detects LLM configuration risks, prompt-injection bait, and agent
structure issues statically — no API keys or network access needed:

```bash
# GenAI-tuned weights: genai findings penalize 35x, handoff only 5x
reproguard scan . --profile genai
```

Real output from scanning a LangGraph + FastAPI agent repository:

```
ReproGuard score: 25/100 (not_ready)
Files scanned: 29 | Issues: 43
Points deducted by category: genai: -35, handoff: -0, privacy: -10, reproducibility: -30

 Severity   Code     Issue                       Location
 ────────   ────     ─────                       ────────
 HIGH       AGENT004 Tool accepts free-form      src/tools/research_tool.py:33
                     input without an allowlist
 HIGH       REPO001  '.env' file is present      .env
                     and not gitignored
 LOW        AGENT002 Agentic graph without       src/agents/base_agent.py:74
                     explicit recursion limit
 LOW        LLMC003  Model ID is not version-    src/config/settings.py:19
                     pinned (gpt-4o-mini)
 LOW        LLMC004  High temperature on         src/config/settings.py:24
                     agentic model (0.7)
```

Run `--format json` for machine-readable findings, or `--model-card` for a
Markdown handoff summarizing risks and recommended actions.

### Custom rules

Add org-specific checks in `.reproguard.yml` — they behave like built-ins
(scored, reported, suppressible via `checks.disabled`/`checks.ignore`):

```yaml
# .reproguard.yml
rules:
  - code: NOSAMPLE      # Flag sampling without a fixed random_state
    title: Sample without seed
    pattern: 'sample\('   # plain regex matched against source text
    category: reproducibility
    severity: medium
    confidence: 0.8
    file_patterns:
      - 'src/**/*.py'
      - '*.ipynb'
  - code: TODO001       # Track tech-debt markers
    title: TODO left in code
    pattern: '# TODO'
    severity: low
```

### Generate reports

```bash
# All report formats
reproguard scan . --format all

# JSON only
reproguard scan . --format json

# SARIF for GitHub Code Scanning
reproguard scan . --format sarif

# Custom output directory
reproguard scan . --output-dir scan-reports
```

### CI integration

```bash
# Fail the build if the score is too low
reproguard scan . --fail-under 50
echo $?  # Exit code 1 when score < 50 or any CRITICAL issue
```

### GitHub Action

Scan in CI with the reusable action:

```yaml
steps:
  - uses: actions/checkout@v4
  - uses: actions/setup-python@v5
    with:
      python-version: "3.11"
  - uses: vipulgote1999/ReproGuard/.github/actions/reproguard-scan@v0.3.1
    with:
      target: "."
      fail-under: 50
      profile: genai
```

### Auto-fixes, model cards, git history

```bash
# Safe, idempotent remediations:
#   gitignore    — append .env protection to .gitignore
#   clear-counts — reset notebook execution counts (clean handoff)
#   seed         — inject missing random/numpy seeds into notebooks
reproguard scan . --fix all          # or pick one: --fix seed

# Markdown handoff artifact (risk summary + recommended actions)
reproguard scan . --model-card

# gitleaks-lite: scan git history diffs for committed secrets
reproguard scan . --git-history

# Validate llm.yaml model IDs against the OpenRouter catalog
reproguard scan . --manifest-online

# GenAI-tuned category weights (or 'profile: genai' in .reproguard.yml)
reproguard scan . --profile genai

# Parallel clean execution of notebooks (kernelspec-aware)
reproguard scan . --execute --parallel --max-workers 4
```

The LLM manifest validated by `--manifest-online` lives at the repo root:

```yaml
# llm.yaml
models:
  - id: gpt-4o-2024-08-06     # pinned — never 'gpt-4o' or '*-latest'
    provider: openai
prompts:
  - id: rag
    path: prompts/rag.md      # must exist on disk
```

Prompt files (`prompts/*.md|txt|yaml|json`) are scanned as whole prompts —
injection-bait phrasing, PII, and oversized prompts are flagged.

### Regression detection (baseline diff)

Gate CI on **new** issues while existing debt is paid down gradually:

```bash
# First run: save the baseline
reproguard scan . --format json --output-dir .reproguard

# Later runs: block on new issues only
reproguard scan . --baseline .reproguard/reproguard-report.json --fail-new 0
reproguard scan . --baseline .reproguard/reproguard-report.json --fail-new-critical
```

New and resolved issues are printed in the terminal summary and recorded in the JSON report metadata. Exit codes: `1` when the gate trips, `2` for invalid baselines.

### Scan with privacy disabled

```bash
reproguard scan . --no-privacy
```

### Custom execution timeout

```bash
reproguard scan notebook.ipynb --execute --execution-timeout 300
```

---

## Understanding Reports

### Score interpretation

| Score | Status | Action Required |
|-------|--------|----------------|
| ≥ 75 | `ready_with_caution` | Review minor issues before production |
| 50–74 | `needs_review` | Address significant issues |
| < 50 | `not_ready` | Blocking issues — must fix |
| Any CRITICAL | `not_ready` | Immediate attention required |
| 0 files scanned | `no_files` | No supported files found — check the scan path (exit code 2) |

### Report files

Reports are written to `.reproguard/` by default:

```
.reproguard/
├── reproguard-report.json      # Structured data
├── reproguard-report.html      # Styled HTML report
└── reproguard-report.sarif     # GitHub Code Scanning compatible
```

---

## Configuration

Create a `.reproguard.yml` in your project root:

```yaml
# .reproguard.yml
exclude_paths:
  - "archive/**"
  - "tests/**"
exclude_dirs:
  - "scratch"
fail_under: 50
checks:
  disabled:
    - "LEAK005"     # Disable large tabular output check
    - "PII004"      # Disable base64 image check
```

Configuration is discovered by walking up from the scan path (like git). See [docs/CONFIGURATION.md](https://github.com/vipulgote1999/ReproGuard/blob/main/docs/CONFIGURATION.md) for the full reference.

---

## Pre-commit Hook

```yaml
# .pre-commit-config.yaml
repos:
  - repo: https://github.com/vipulgote1999/ReproGuard
    rev: v0.3.1
    hooks:
      - id: reproguard-scan
        args: ["--fail-under", "75"]
```

The hook scans the entire repository on each commit and blocks the commit when the score falls below the threshold.

---

## CI/CD Integration

### GitHub Actions (with SARIF upload)

```yaml
name: ReproGuard
on: [push, pull_request]
jobs:
  scan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.11"
      - run: pip install reproguard
      - run: reproguard scan . --format sarif --fail-under 50
      - uses: github/codeql-action/upload-sarif@v3
        with:
          sarif_file: .reproguard/reproguard-report.sarif
```

### GitLab CI

```yaml
reproguard:
  stage: test
  script:
    - pip install reproguard
    - reproguard scan . --format html --fail-under 50
  artifacts:
    paths:
      - .reproguard/
```

---

## Project Structure

```
ReproGuard/
├── src/reproguard/
│   ├── cli.py              # Typer CLI entry point
│   ├── scanner.py          # Scan orchestrator
│   ├── models.py           # Core data models & scoring
│   ├── config.py           # .reproguard.yml loader (profiles, rules)
│   ├── python_analysis.py  # Python source analysis
│   ├── notebook.py         # Notebook parser + hidden-state analysis
│   ├── leakage.py          # ML data leakage heuristics
│   ├── privacy.py          # PII / secret scanning
│   ├── dependency.py       # Dependency file analysis
│   ├── execution.py        # Clean execution (parallel, kernel-aware)
│   ├── repo.py             # Repository hygiene checks (REPO001-007)
│   ├── genai.py            # LLM config / prompt / agent checks
│   ├── unsafe.py           # Bandit-lite unsafe Python checks
│   ├── rules.py            # Custom rule engine
│   ├── manifest.py         # llm.yaml manifest validation
│   ├── fixes.py            # --fix auto-remediation
│   ├── githistory.py       # --git-history secret scan
│   ├── modelcard.py        # --model-card generation
│   ├── report.py           # JSON/HTML report generators
│   ├── sarif.py            # SARIF 2.1.0 report generator
│   ├── plugin.py           # Check registry & filtering
│   └── utils.py            # Shared helpers
├── examples/
│   ├── risky_customer_churn.ipynb  # Notebook with intentional issues
│   ├── clean_analysis.py           # Clean script example
│   └── requirements.txt            # Example dependency file
├── docs/
│   ├── ARCHITECTURE.md     # Internal design & module docs
│   ├── CHECKS.md           # Complete issue code reference
│   ├── CONFIGURATION.md    # Configuration file reference
│   ├── GUIDES.md           # Usage guides & integrations
│   └── ROADMAP.md          # Future plans
├── pyproject.toml          # Build & tool config
└── README.md               # This file
```

---

## Check Reference

| Code | Check | Severity | Category |
|------|-------|----------|----------|
| AGENT001 | Graph Without Checkpointing | ? | genai |
| AGENT002 | No Recursion Limit | ? | genai |
| AGENT003 | Tool Without Docstring | ? | genai |
| AGENT004 | Tool With Unvalidated Input | ? | genai |
| AGENT005 | Privileged Tool Without Approval | ? | genai |
| AGENT006 | Tool Executes Shell or Dynamic Code | ? | genai |
| AGENT007 | Incomplete CrewAI Agent | ? | genai |
| DATA001 | Local Machine Path | ? | data_dependency |
| DATA002 | Missing Data File | ? | data_dependency |
| DEP001 | No Dependency File | ? | reproducibility |
| DEP002 | Unpinned Dependency | ? | reproducibility |
| DEP003 | Missing Imported Package | ? | reproducibility |
| EXEC001 | Notebook Execution Error | ? | execution |
| EXEC002 | Execution Setup Failure | ? | execution |
| GEN001 | System Prompt Leakage | ? | genai |
| GEN002 | LLM Output Executed | ? | genai |
| GEN003 | Test Data in Training | ? | genai |
| HAND001 | Missing Handoff Documentation | ? | handoff |
| LEAK001 | Preprocessing Before Split | ? | data_leakage |
| LEAK002 | Future/Target Column Name | ? | data_leakage |
| LEAK003 | Test Data in Fit Call | ? | data_leakage |
| LEAK004 | Suspiciously High Metric | ? | data_leakage |
| LEAK005 | Large Tabular Output | ? | data_leakage |
| LEAK006 | Hardcoded Path in Output | ? | data_dependency |
| LEAK007 | Exception Traceback in Output | ? | reproducibility |
| LLMC001 | LLM Client Without Temperature | ? | genai |
| LLMC002 | LLM Call Without Output Limits | ? | genai |
| LLMC003 | Unpinned Model Identifier | ? | genai |
| LLMC004 | High Temperature on Agentic Model | ? | genai |
| LLMC005 | LLM Client Without Timeout or Retries | ? | genai |
| LLMC006 | Vector Search Without Top-K | ? | genai |
| LLMC007 | Text Splitter Without Chunk Size | ? | genai |
| LLMC008 | Hardcoded API Base URL | ? | genai |
| MAN001 | Manifest Model Not Pinned | ? | genai |
| MAN002 | Manifest Prompt Path Missing | ? | genai |
| MAN003 | Unknown Manifest Provider | ? | genai |
| MAN004 | Unparseable LLM Manifest | ? | genai |
| NB001 | Out-of-Order Execution | ? | reproducibility |
| NB002 | Stale Output Without Execution Count | ? | reproducibility |
| NB003 | Non-Default Kernel Requirement | ? | reproducibility |
| NB004 | Use Before Assign Across Cells | ? | reproducibility |
| NB005 | Undefined Name | ? | reproducibility |
| NB006 | No Markdown Cells | ? | reproducibility |
| PII001 | Email Address | ? | privacy |
| PII002 | Phone Number | ? | privacy |
| PII003 | Credit Card Number | ? | privacy |
| PII004 | Base64 Image in Output | ? | privacy |
| PII005 | Large JSON Blob | ? | privacy |
| PROMPT001 | Missing System Prompt | ? | genai |
| PROMPT002 | Prompt Injection Bait | ? | genai |
| PROMPT003 | Untrusted Content in Prompt | ? | genai |
| PROMPT004 | Oversized Prompt | ? | genai |
| PROMPT005 | Duplicated Prompt | ? | genai |
| AGENT001 | Graph Without Checkpointing | MEDIUM | genai |
| AGENT002 | No Recursion Limit | LOW | genai |
| AGENT003 | Tool Without Docstring | LOW | genai |
| AGENT004 | Tool With Unvalidated Input | HIGH | genai |
| AGENT005 | Privileged Tool Without Approval | HIGH | genai |
| AGENT006 | Tool Executes Shell or Dynamic Code | HIGH | genai |
| AGENT007 | Incomplete CrewAI Agent | LOW | genai |
| DATA001 | Local Machine Path | HIGH | data_dependency |
| DATA002 | Missing Data File | MEDIUM | data_dependency |
| DEP001 | No Dependency File | HIGH | reproducibility |
| DEP002 | Unpinned Dependency | LOW | reproducibility |
| DEP003 | Missing Imported Package | MEDIUM | reproducibility |
| EXEC001 | Notebook Execution Error | CRITICAL | execution |
| EXEC002 | Execution Setup Failure | HIGH | execution |
| GEN001 | System Prompt Leakage | MEDIUM | genai |
| GEN002 | LLM Output Executed | CRITICAL | genai |
| GEN003 | Test Data in Training | HIGH | genai |
| HAND001 | Missing Handoff Documentation | LOW | handoff |
| LEAK001 | Preprocessing Before Split | CRITICAL | data_leakage |
| LEAK002 | Future/Target Column Name | HIGH | data_leakage |
| LEAK003 | Test Data in Fit Call | CRITICAL | data_leakage |
| LEAK004 | Suspiciously High Metric | MEDIUM | data_leakage |
| LEAK005 | Large Tabular Output | LOW | data_leakage |
| LEAK006 | Hardcoded Path in Output | MEDIUM | data_dependency |
| LEAK007 | Exception Traceback in Output | HIGH | reproducibility |
| LLMC001 | LLM Client Without Temperature | MEDIUM | genai |
| LLMC002 | LLM Call Without Output Limits | MEDIUM | genai |
| LLMC003 | Unpinned Model Identifier | LOW | genai |
| LLMC004 | High Temperature on Agentic Model | MEDIUM | genai |
| LLMC005 | LLM Client Without Timeout or Retries | LOW | genai |
| LLMC006 | Vector Search Without Top-K | LOW | genai |
| LLMC007 | Text Splitter Without Chunk Size | LOW | genai |
| LLMC008 | Hardcoded API Base URL | LOW | genai |
| MAN001 | Manifest Model Not Pinned | LOW | genai |
| MAN002 | Manifest Prompt Path Missing | MEDIUM | genai |
| MAN003 | Unknown Manifest Provider | LOW | genai |
| MAN004 | Unparseable LLM Manifest | MEDIUM | genai |
| NB001 | Out-of-Order Execution | MEDIUM | reproducibility |
| NB002 | Stale Output Without Execution Count | LOW | reproducibility |
| NB003 | Non-Default Kernel Requirement | MEDIUM | reproducibility |
| NB004 | Use Before Assign Across Cells | HIGH | reproducibility |
| NB005 | Undefined Name | HIGH | reproducibility |
| NB006 | No Markdown Cells | LOW | reproducibility |
| PII001 | Email Address | HIGH | privacy |
| PII002 | Phone Number | MEDIUM | privacy |
| PII003 | Credit Card Number | HIGH | privacy |
| PII004 | Base64 Image in Output | LOW | privacy |
| PII005 | Large JSON Blob | LOW | privacy |
| PROMPT001 | Missing System Prompt | MEDIUM | genai |
| PROMPT002 | Prompt Injection Bait | HIGH | genai |
| PROMPT003 | Untrusted Content in Prompt | HIGH | genai |
| PROMPT004 | Oversized Prompt | MEDIUM | genai |
| PROMPT005 | Duplicated Prompt | LOW | genai |
| PY001 | Python Syntax Error | CRITICAL | reproducibility |
| REP001 | Missing Random Seed | MEDIUM | reproducibility |
| REPO001 | Committed .env File | HIGH | privacy |
| REPO002 | Missing Environment Template | LOW | reproducibility |
| REPO003 | Large Committed Model Artifact | MEDIUM | data_dependency |
| REPO004 | Unversioned or Empty Data Directory | MEDIUM | data_dependency |
| REPO005 | Tests Without CI | LOW | reproducibility |
| REPO006 | Missing or Minimal README | LOW | handoff |
| REPO007 | DVC References Missing Data | MEDIUM | data_dependency |
| SEC001 | AWS Access Key | CRITICAL | privacy |
| SEC002 | Secret Assignment | CRITICAL | privacy |
| SEC003 | Private Key Block | CRITICAL | privacy |
| SEC004 | High-Entropy String | MEDIUM | privacy |
| SEC005 | Dynamic Code Execution | HIGH | privacy |
| SEC006 | Unsafe YAML Loading | HIGH | privacy |
| SEC007 | Shell Command Execution | HIGH | privacy |
| SEC008 | Unpickle of Untrusted Input | MEDIUM | privacy |
| SEC009 | HTTP Call Without Timeout | LOW | privacy |
| SEC010 | SQL String Interpolation | MEDIUM | privacy |
| SEC011 | Assert Outside Tests | INFO | privacy |

See [docs/CHECKS.md](https://github.com/vipulgote1999/ReproGuard/blob/main/docs/CHECKS.md) for full details on every check.

---

## Design Principles

1. **Local-first** — Scans run entirely on your machine. No data or notebooks leave your environment.
2. **Explainable rules** — Every issue has a code, severity, evidence, confidence score, and suggested fix. No black boxes.
3. **Low friction** — CLI-first design with a single `reproguard scan <target>` command. Pre-commit hook, CI integration, and GitHub Action out of the box.
4. **Conservative scoring** — The score is transparent (penalty-based, weighted by severity and category). You can customize all penalties and weights.
5. **Narrow wedge** — Focused on catching pre-production risks before work enters heavier MLOps pipelines.

---

## Limitations

ReproGuard v0.3 (alpha) uses heuristics. It flags likely risks but cannot prove every issue is real. Treat it as a **review assistant**, not a final governance decision.

- Leakage detection is heuristic — expect false positives (use `checks.ignore` to silence known-safe locations)
- Dependency parsing is intentionally lightweight (regex-based for requirements/Pipfile, YAML for conda env files, TOML for pyproject/lock files — no full resolver)
- Privacy scanning uses regex rules by default (Presidio support planned)
- Clean notebook execution depends on local kernel and dependency availability
- Data file existence checks are limited to paths referenced from notebooks/scripts

---

## Development

```bash
# Install dev dependencies
pip install -e .[dev]

# Lint
ruff check .

# Test
pytest

# Build release artifacts and validate metadata
python -m build
python -m twine check dist/*
```

Releases follow the procedure in [docs/RELEASING.md](https://github.com/vipulgote1999/ReproGuard/blob/main/docs/RELEASING.md) — see
the checklist there before tagging. Changes are tracked in
[CHANGELOG.md](https://github.com/vipulgote1999/ReproGuard/blob/main/CHANGELOG.md).

---

## Roadmap

- **v0.2** ✅ — correctness hardening (magic handling, conda/lock dependency
  parsing), path-scoped ignore rules, baseline diffing, extended coverage
- **v0.3** ✅ — GenAI category (27 checks), repository hygiene, custom rule
  engine, profiles, auto-fixes, git-history secret scan, LLM manifests, model
  cards, parallel + kernel-aware execution, GitHub Action
- **v0.4 (planned)**: prompt-suite drift detection, `--fix` for more codes,
  SARIF-to-CodeQL grouping, notebook diffs, license scanning
- **v1.0+**: ML pipeline scanning, differential scans, team dashboard, API

See [docs/ROADMAP.md](https://github.com/vipulgote1999/ReproGuard/blob/main/docs/ROADMAP.md) for the full roadmap.

---

## License & Usage Rights

© 2026 **Vipul Gote**. All rights reserved.

ReproGuard is **free to use** for its intended purpose — pre-production risk
scanning of data science notebooks, Python scripts, and ML repositories — but
it is **not open source**. You may use the tool as-is for your own scanning
work, but you may **not**:

- copy, reproduce, or clone the source code beyond what is needed to run it;
- modify or build derivative works from it;
- redistribute, sell, sublicense, or offer it as a hosted service;
- reuse its code, heuristics, or ideas to build a competing tool;
- claim authorship or remove the copyright notice.

All other use requires prior written permission from Vipul Gote. See
[LICENSE](https://github.com/vipulgote1999/ReproGuard/blob/main/LICENSE) for the complete terms.
