Metadata-Version: 2.4
Name: releaseguard-cli
Version: 0.1.0
Summary: Scan a dataset or model directory for PII and secrets, redact what you find, and package a public-release bundle with a Hugging Face dataset/model card and an EU AI Act Art. 53(1)(d) training-data summary -- generated from the scan, not hand-written.
Project-URL: Homepage, https://github.com/RudrenduPaul/ReleaseGuard
Project-URL: Repository, https://github.com/RudrenduPaul/ReleaseGuard
Project-URL: Issues, https://github.com/RudrenduPaul/ReleaseGuard/issues
Project-URL: Changelog, https://github.com/RudrenduPaul/ReleaseGuard/blob/main/CHANGELOG.md
Project-URL: Author - Rudrendu Paul, https://github.com/RudrenduPaul
Project-URL: Author - Sourav Nandy, https://github.com/Sourav-nandy-ai
Author: Rudrendu Paul, Sourav Nandy
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: agent-tools,ai-compliance,data-redaction,dataset-card,eu-ai-act,mcp,model-card,pii-detection,presidio,responsible-ai
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Security
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Requires-Dist: click>=8.1
Requires-Dist: presidio-analyzer>=2.2.360
Requires-Dist: presidio-anonymizer>=2.2.360
Requires-Dist: pyyaml>=6.0
Requires-Dist: spacy>=3.7
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == 'dev'
Requires-Dist: mypy>=1.10; extra == 'dev'
Requires-Dist: pytest-cov>=5.0; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Requires-Dist: types-pyyaml>=6.0; extra == 'dev'
Provides-Extra: mcp
Requires-Dist: mcp>=2.0; extra == 'mcp'
Provides-Extra: parquet
Requires-Dist: pyarrow>=15.0; extra == 'parquet'
Description-Content-Type: text/markdown

# ReleaseGuard

**Scan a dataset or model directory for PII and secrets with [Presidio](https://github.com/data-privacy-stack/presidio), redact what you find, and generate a public-release bundle, a Hugging Face dataset/model card plus an EU AI Act Art. 53(1)(d) training-data summary, in one command.**

[![CI](https://github.com/RudrenduPaul/ReleaseGuard/actions/workflows/ci.yml/badge.svg)](https://github.com/RudrenduPaul/ReleaseGuard/actions/workflows/ci.yml)
[![License: Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](LICENSE)
[![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue)](pyproject.toml)

<!-- TODO: record a real terminal-recording demo GIF (vhs) once the CLI is on PyPI; until then the Quickstart section below is a real, unedited terminal transcript, not a mockup. -->

ReleaseGuard is **not** a PII detector. It's the glue between "I have a dataset I want to publish" and "I have a sanitized bundle with the paperwork already drafted." Detection is entirely [Presidio](https://github.com/data-privacy-stack/presidio)'s, an actively maintained open-source project with 10,000+ GitHub stars. ReleaseGuard chains Presidio's scan straight into redaction and into the two documents almost every public dataset/model release actually needs, instead of you writing a script to do it yourself.

## Table of contents

- [Quick summary](#quick-summary)
- [Install](#install)
- [Quickstart](#quickstart)
- [CLI reference](#cli-reference)
- [Library API](#library-api)
- [Agent-native (MCP + A2A)](#agent-native-mcp--a2a)
- [Comparison](#comparison)
- [What is ReleaseGuard, and why does it exist](#what-is-releaseguard-and-why-does-it-exist)
- [What ReleaseGuard is not](#what-releaseguard-is-not)
- [FAQ](#faq)
- [Contributing](#contributing)

## Quick summary

- **Install:** `pip install releaseguard-cli` then `python -m spacy download en_core_web_sm` (Presidio's NLP model, one-time, ~13 MB)
- **Use it for:** turning a Presidio scan into a redacted copy plus a dataset/model card and an EU AI Act Art. 53(1)(d) training-data summary, from one command instead of three separate tools and a hand-written template
- **What it's not:** a PII detector of its own, or a claim of "total transparency"/full training-data disclosure compliance. See [What ReleaseGuard is not](#what-releaseguard-is-not)
- **Runs entirely local.** No dataset content, scan results, or redacted output is ever sent to a remote service.

## Install

```bash
pip install releaseguard-cli
python -m spacy download en_core_web_sm   # Presidio's default NLP model (~13 MB, one-time)

# npm launcher (thin wrapper around the PyPI package, see "Why two registries" in FAQ)
npx releaseguard-cli --help
```

`en_core_web_sm` is spaCy's small English model and Presidio's own quickstart default. For higher-accuracy detection, install the larger model instead and pass `--spacy-model`:

```bash
python -m spacy download en_core_web_lg   # ~400 MB, Presidio's recommendation for production
releaseguard scan ./data --spacy-model en_core_web_lg
```

## Quickstart

This is a real, unedited run against a two-row sample CSV, not a mockup:

```
$ releaseguard scan dataset --score-threshold 0.4
Scanned 1 file(s) under dataset
Total findings: 8

entity type                 count
URL                         3
PERSON                      2
EMAIL_ADDRESS               2
ORGANIZATION                1
```

That `URL: 3` line is a real Presidio quirk worth calling out rather than hiding: its regex-based URL recognizer also fires on the domain portion of an email address (`example.com` inside `alice.rivera@example.com`), so email-heavy text double-counts as both `EMAIL_ADDRESS` and `URL`. This is Presidio's own recognizer behavior, unmodified. Filter it out with `--entities` if you only care about email addresses:

```bash
releaseguard scan dataset --entities EMAIL_ADDRESS,PERSON,PHONE_NUMBER
```

Redact, then package a release bundle in one command:

```bash
$ releaseguard package dataset --output bundle --score-threshold 0.4
Release bundle written to bundle
  dataset card: bundle/README-dataset-card.md
  EU AI Act Art. 53(1)(d) summary: bundle/eu-ai-act-training-summary.md
  redacted source: bundle-redacted-source
```

`bundle/eu-ai-act-training-summary.md` opens like this, real scan counts filled in, everything else left as an explicit placeholder for a human to complete:

```markdown
# Training Data Summary (EU AI Act Art. 53(1)(d))

_Draft generated by ReleaseGuard... Based on the European Commission's
training-data-summary template, published 2025-07-24. This covers the
narrow, currently-binding categorical-summary requirement only -- it is
not a claim of full training-data disclosure._

## 4. Personal Data and PII Handling (from ReleaseGuard scan)

A Presidio-backed scan (detector: `presidio`) covered 1 file(s) under `dataset`.

| PII/secret category detected | Occurrences |
| --- | --- |
| `EMAIL_ADDRESS` | 2 |
| `ORGANIZATION` | 1 |
| `PERSON` | 2 |
| `URL` | 3 |
```

`--json` on every command switches to machine-readable output for scripts and agents.

## CLI reference

```
$ releaseguard --help
Usage: releaseguard [OPTIONS] COMMAND [ARGS]...

  Scan, redact, and package a dataset/model directory for public release.

Options:
  --version  Show the version and exit.
  --help     Show this message and exit.

Commands:
  mcp      Start an MCP server exposing scan/redact/package as agent tools.
  package  Scan PATH, optionally redact it, and generate a release bundle...
  redact   Scan PATH and write a redacted copy to --output.
  scan     Scan PATH (a file or directory) for PII and secrets using...
```

| Command | Purpose |
| --- | --- |
| `scan PATH` | Scans a file or directory (CSV, JSON/JSONL, plain text) with Presidio. `--entities`, `--score-threshold`, `--spacy-model`, `--json`. |
| `redact PATH --output DIR` | Scans, then writes a redacted copy to `DIR`. Never touches `PATH`. `--strategy mask\|hash\|remove`, `--overwrite`. |
| `package PATH --output DIR` | Scans (and by default redacts first), then writes a Hugging Face dataset/model card plus the EU AI Act Art. 53(1)(d) summary to `DIR`. `--kind dataset\|model\|both`, `--redact-first/--no-redact-first`. |
| `mcp` | Starts an MCP server (stdio) exposing `scan_directory_tool`, `redact_directory_tool`, `package_release_tool`. Requires `pip install "releaseguard-cli[mcp]"` on Python 3.10+. |

Every command supports `--json`. Full flag reference: `releaseguard <command> --help`.

## Library API

```python
from releaseguard.detectors import get_detector
from releaseguard.scanner import scan_directory
from releaseguard.redactor import redact_directory
from releaseguard.packager import build_release_bundle
from releaseguard.types import RedactionStrategy

detector = get_detector("presidio", score_threshold=0.4)
scan_result = scan_directory("dataset/", detector)

redaction_result = redact_directory(
    scan_result, "dataset-redacted/", strategy=RedactionStrategy.MASK
)

bundle = build_release_bundle(
    scan_result, "bundle/", redaction_result=redaction_result, source_kind="dataset"
)
print(bundle.eu_ai_act_summary_path)
```

`PIIDetector` (`releaseguard.detectors.base`) and `FileReader` (`releaseguard.readers.base`) are the two extension points. Presidio is the only detector shipped in v0.1; CSV, JSON/JSONL, and plain text are the three readers shipped in v0.1. Both are registries, not hardcoded calls, specifically so a new format or a second detection backend is a scoped addition later. See [CONTRIBUTING.md](CONTRIBUTING.md).

## Agent-native (MCP + A2A)

```bash
pip install "releaseguard-cli[mcp]"
releaseguard mcp
```

Exposes three tools over stdio MCP: `scan_directory_tool`, `redact_directory_tool`, `package_release_tool`, each returning the same JSON shape as the matching CLI `--json` output. A `.well-known/agent.json` manifest is shipped at the repo root for A2A-style discovery, listing both the CLI and MCP interfaces and the packages that provide them.

## Comparison

| | ReleaseGuard | Presidio | huggingface_hub card tooling | `pii-lib` |
|---|---|---|---|---|
| Detects PII/secrets | No, wraps Presidio | **Yes** (its own job) | No | Yes (regex + NER, code-training-data scoped) |
| Redacts detected entities | Yes (via `presidio-anonymizer`) | Yes (library-level) | No | Yes |
| Generates a Hugging Face dataset/model card | **Yes**, from real scan results | No | Yes (manual fields, no scan integration) | No |
| Generates an EU AI Act Art. 53(1)(d) summary | **Yes**, from real scan results | No | No | No |
| One command, scan through release bundle | **Yes** | No (library only, you write the glue) | No (library only) | No |
| GitHub stars (checked 2026-08) | New in 2026 | ~10,300 | ~3,800 | 16 |

Star counts checked live against the GitHub API on 2026-08-03: [data-privacy-stack/presidio](https://github.com/data-privacy-stack/presidio) (originally `microsoft/presidio`; the project moved organizations, same codebase), [huggingface/huggingface_hub](https://github.com/huggingface/huggingface_hub), [bigcode-project/pii-lib](https://github.com/bigcode-project/pii-lib). `pii-lib`'s low star count is itself informative: it is the closest prior attempt at PII redaction scoped to a training-data release workflow, and it has not gained meaningful adoption. ReleaseGuard does not assume that outcome will be different here; see the FAQ entry on demand.

Enterprise data-governance platforms (Databricks Unity Catalog, Credo AI, BigID, and others) already offer PII classification and redaction as part of broader paid platforms aimed at large organizations. ReleaseGuard is a free, single-purpose, open-source alternative for a team that just wants the scan-redact-package workflow for one release, not a governance suite.

## What is ReleaseGuard, and why does it exist

ReleaseGuard is an open-source CLI, Python library, and MCP server that chains three steps, PII/secret detection (via Presidio), redaction, and public-release documentation, into one command. Each step already exists as a separate tool: Presidio detects, `presidio-anonymizer` redacts, `huggingface_hub` has card-generation helpers, and the European Commission publishes a training-data-summary template as a document you fill in by hand. Nothing before ReleaseGuard chained a real scan directly into a filled-in template.

It exists because publishing a dataset or model responsibly involves running a PII scan, redacting what it finds, and then writing up two documents almost by hand, a card and (for general-purpose AI model providers) an EU AI Act training-data summary. ReleaseGuard automates the second half of that workflow so the resulting documents reflect what was actually scanned, not what someone remembered to write down afterward.

**EU AI Act Art. 53(1)(d), stated precisely:** this article requires providers of *general-purpose AI (GPAI) models* to publish a "sufficiently detailed summary" of training content, using the template the European Commission's AI Office published on 2025-07-24 (in force for new models from 2025-08-02, transitional deadline 2027-08-02 for models already on the market, enforcement checks from the AI Office starting 2026-08-02). It requires a **categorical summary of data sources and modalities**. It does not require raw training samples, full training recipes, or model weights, and it applies specifically to GPAI model providers, not to every dataset publisher. ReleaseGuard's generated summary is a starting draft for that narrow, real requirement, never a claim of broader "total transparency" compliance.

## What ReleaseGuard is not

- **Not a PII detector.** Every entity type, confidence score, and detection decision comes from Presidio. ReleaseGuard adds no NLP model, no recognizer, and no accuracy claim of its own. If Presidio misses something or overcounts (see the `URL`/email overlap in the Quickstart above), ReleaseGuard inherits that behavior unmodified.
- **Not a hosted service.** Everything runs locally. No scan target, scan result, or redacted output is transmitted anywhere. See [SECURITY.md](SECURITY.md)'s scope section.
- **Not proof of legal compliance.** The generated EU AI Act summary is a draft that still needs a human to fill in licensing, data-source, and copyright sections ReleaseGuard cannot infer from a scan. Running `releaseguard package` does not, by itself, satisfy Art. 53(1)(d) or any other regulation.
- **Not evidence of demand beyond what's cited above.** The closest prior attempt at this same workflow, `pii-lib`, sits at 16 GitHub stars. ReleaseGuard does not claim to have solved the adoption problem that project ran into; it claims to fill a real, narrow, independently-verified gap (no existing open-source tool chains a Presidio scan directly into an Art. 53(1)(d) template), and lets real usage decide the rest.

## FAQ

**Does ReleaseGuard detect PII more accurately than Presidio?**
No. It cannot, since it calls Presidio's own `AnalyzerEngine` for every detection decision. Any accuracy question is a Presidio question; see [Presidio's own documentation](https://github.com/data-privacy-stack/presidio) and [presidio-research](https://github.com/microsoft/presidio-research) for its evaluation methodology.

**Why does `scan` need a spaCy model download?**
Presidio's `AnalyzerEngine` requires a spaCy language model for context-aware detection (recognizing that "John Smith" is a name from surrounding text, not just a capitalized word). spaCy models ship as their own installable packages, not as a `pip` dependency, so `python -m spacy download en_core_web_sm` is a required one-time step, the same as it is for anyone using Presidio directly.

**Is the EU AI Act Art. 53(1)(d) summary legally sufficient on its own?**
No. It is a structurally correct starting draft populated with real scan data where ReleaseGuard can verify it (the PII/secrets section) and an explicit placeholder everywhere it can't (data sources, licensing, copyright status). A human, ideally with legal review, has to fill in the placeholders before publishing it as a compliance artifact.

**Does this only apply if I'm training a GPAI model?**
The Art. 53(1)(d) summary specifically targets general-purpose AI model providers under the EU AI Act, a narrow buyer set. The `scan`, `redact`, and Hugging Face card-generation parts of ReleaseGuard are useful for any dataset or model release, regardless of whether Art. 53(1)(d) applies to you.

**Why two registries?**
ReleaseGuard's implementation is Python, since Presidio itself is Python (`presidio-analyzer`/`presidio-anonymizer`); wrapping it in another language would mean re-shelling out or reimplementing bindings. The npm package (`releaseguard-cli`) is a thin launcher, not a reimplementation. It locates and execs the real `releaseguard` binary installed from PyPI, so `npx releaseguard-cli` works for npm-first agent tooling without duplicating Presidio's detection logic in two languages.

**Does anyone actually need this, or is it "glue code nobody asked for"?**
Honestly stated: no organic demand signal (an HN/Reddit thread describing this exact workflow as a lived pain point) had surfaced as of this project's initial research. The independently verifiable fact is narrower and more defensible: no existing open-source tool chains a Presidio scan directly into an Art. 53(1)(d) template or a combined HF card, in one command, from one scan. Whether that gap turns into real usage is an open, falsifiable question this project tracks rather than assumes the answer to.

**What happens to files ReleaseGuard can't parse (images, model weight files, Parquet without the extra)?**
`scan` and `redact` skip them (listed under `files_skipped` in `--json` output); `redact` copies them through to the output directory unchanged rather than silently dropping them from the release bundle. They are not scanned for PII, so review them separately before publishing.

## Contributing

See [CONTRIBUTING.md](CONTRIBUTING.md). Security issues: see [SECURITY.md](SECURITY.md).

## License

[Apache 2.0](LICENSE)
