Metadata-Version: 2.4
Name: pseudopros
Version: 0.1.1
Summary: Cryptographically secure facial anonymization for computer vision pipelines
Project-URL: Homepage, https://github.com/dactylroot/pseudopros
Project-URL: Repository, https://github.com/dactylroot/pseudopros
Project-URL: Issues, https://github.com/dactylroot/pseudopros/issues
Author-email: Cory Root <dactylroot@gmail.com>
License: MIT License
        
        Copyright (c) 2024 Cory Root
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
License-File: LICENSE
Keywords: anonymization,computer-vision,cryptography,face,privacy,re-identification,video
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Scientific/Engineering :: Image Processing
Classifier: Topic :: Security :: Cryptography
Requires-Python: >=3.12
Requires-Dist: accelerate>=0.30
Requires-Dist: cryptography>=41.0
Requires-Dist: diffusers<0.31,>=0.28
Requires-Dist: huggingface-hub>=0.20
Requires-Dist: imageio-ffmpeg>=0.4
Requires-Dist: imageio>=2.31
Requires-Dist: insightface>=0.7
Requires-Dist: numpy<2.0,>=1.24
Requires-Dist: onnxruntime<1.20,>=1.16
Requires-Dist: opencv-python>=4.8
Requires-Dist: pillow>=10.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: scikit-image<0.24,>=0.21
Requires-Dist: scipy<1.14,>=1.11
Requires-Dist: torch<2.3,>=2.2
Requires-Dist: torchvision<0.18,>=0.17
Provides-Extra: dev
Requires-Dist: hatch>=1.12; extra == 'dev'
Requires-Dist: pytest-cov>=5.0; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: rtsp>=2.0.2; extra == 'dev'
Provides-Extra: secondary
Requires-Dist: mediapipe<0.10.30,>=0.10; extra == 'secondary'
Description-Content-Type: text/markdown

# Pseudopros

Pseudo (false) + Prosopon (face/mask)

    ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢀⣠⣾⡆⠀⠀⠀⠀
    ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣠⣶⣿⣿⣿⠃⣠⠀⠀⠀
    ⠀⠀⠀⣄⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⢀⣀⣤⣶⣿⣿⣿⣿⣿⠇⢰⣿⣇⠀⠀
    ⠀⠀⣼⣿⣿⣶⣤⣀⡀⠀⠀⠀⠀⠻⠿⠟⠛⣿⣿⣿⣿⣿⣿⡏⠀⠛⣿⣿⡄⠀
    ⠀⢠⣿⣿⣿⣿⣿⣿⣿⣿⣿⣶⣶⣶⣶⡆⢀⡛⠛⠛⠛⢿⣿⣆⣀⣠⣿⣿⡧⠀
    ⠀⢸⣉⠁⠀⢠⣿⣿⠁⣈⡙⠛⠻⠿⠿⠃⠸⠿⠇⠀⠀⢈⡁⢸⣿⣿⡿⠟⣡⠀
    ⠀⣿⣿⣆⣀⣸⣿⡿⠀⣿⣿⣿⣷⣶⣶⣶⣶⣶⣶⣾⣿⣿⣇⠈⠟⢉⣤⣾⡇⠀
    ⠀⢸⣿⣿⣿⣿⣿⡇⠀⠿⠿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⡿⠿⠀⣶⣿⣿⠟⠀⠀
    ⠀⠈⣉⣉⣉⣉⡉⠓⠀⣶⣦⣤⠀⠉⠙⣿⣿⡏⠉⠀⣤⣤⡖⠀⠉⠉⠁⠀⠀⠀
    ⠀⠀⠘⢿⣿⣿⣿⣿⡀⢿⣿⣿⣶⣤⠀⣿⣿⠁⢠⣶⣿⣿⠇⠀⠀⠀⠀⠀⠀⠀
    ⠀⠀⠀⠈⠛⠻⠿⠟⠃⠘⣿⣿⣿⣿⣶⣿⣿⣶⣿⣿⣿⡿⠀⠀⠀⠀⠀⠀⠀⠀
    ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠹⣿⣯⣉⠉⠉⠉⢉⣩⣿⡿⠁⠀⠀⠀⠀⠀⠀⠀⠀
    ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠘⢿⣿⣿⣶⣿⣿⣿⠟⠁⠀⠀⠀⠀⠀⠀⠀⠀⠀
    ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠈⠙⠛⠛⠉⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀

A Python library for cryptographically secure facial anonymization in computer vision pipelines.

Pseudopros replaces real faces in video streams with deterministic synthetic faces that:

- **Cannot be re-identified** - output faces are cryptographically unlinked from the original (validated against ArcFace ReID)
- **Are repeatable** - the same input face always produces the same synthetic output for a given key and salt
- **Are cross-source** - the same key and salt produce consistent identities across independently captured video of the same subject
- **Preserve motion** - expression, pose, and movement data are retained for downstream tracking and aggregate analysis

Intended for use within streaming pipelines (e.g. GStreamer). Keys and salts are passed as parameters and should be stored remotely, never embedded with the video.

## Security properties

| Property | Mechanism |
|---|---|
| Re-identification resistance | LDM-generated face is orthogonal to original in ArcFace space (validated: similarity ≈ -0.03 vs threshold 0.28) |
| Repeatability | HKDF is deterministic; same inputs always produce the same seed and therefore the same synthetic face |
| Cross-source consistency | Key derivation depends only on the identity embedding, not the video source |
| Key isolation | Rotating the salt changes all synthetic identities without changing the secret |

## Cryptography

pseudopros depends on the [PyCA `cryptography`](https://cryptography.io) library for one purpose: **HKDF-SHA256** key derivation in `keying.py`.

**HKDF** (HMAC-based Key Derivation Function, [RFC 5869](https://datatracker.ietf.org/doc/html/rfc5869)) turns a secret key and a subject embedding into a fixed-length seed:

| Input | Value |
|---|---|
| IKM (input key material) | Your secret bytes (`secret`) |
| Salt | Per-deployment salt (`salt`) |
| Info | `b"pseudopros-v1:" + subject_embedding_bytes` |
| Output length | 32 bytes |

The `info` field provides domain separation: it binds the derived seed to this library version and to the specific subject identity. Changing the subject embedding, the salt, or the secret each produces an independent seed.

No encryption, signing, MAC verification, or other crypto primitives are used anywhere else in the codebase.

## How it works

1. **Enrollment** - ArcFace embeds the subject's face into a 512-dimensional identity vector
2. **Key derivation** - HKDF-SHA256 over `(secret, salt, identity_vector)` produces a 32-byte seed
3. **Synthetic face generation** - the seed drives a Latent Diffusion Model to produce a deterministic synthetic face unrelated to the original
4. **Animation** - First Order Motion Model transfers the subject's motion onto the synthetic face frame by frame
5. **Paste-back** - the animated region is blended back into the original frame with a soft oval mask

The synthetic face is the cryptographic output: it is a pure function of `(secret, salt, subject)` with no recoverable path to the original identity.

## Installation

```bash
pip install pseudopros
```

Requires Python 3.12. PyTorch is installed from the PyTorch CPU index on Intel Mac:

```bash
pip install pseudopros --extra-index-url https://download.pytorch.org/whl/cpu
```

On first use, pseudopros downloads three models:

- **InsightFace buffalo_l** (~280 MB, cached in `~/.insightface/`) - ArcFace face detection and embedding
- **dactylroot/pseudopros-fomm** (~695 MB, cached in `~/.cache/huggingface/`) - First Order Motion Model checkpoint
- **CompVis/ldm-celebahq-256** (~1.3 GB, cached in `~/.cache/huggingface/`) - Latent Diffusion Model for synthetic face generation (only when using `RosterSynthesizer`)

## Usage

pseudopros has two deployment modes, reflected directly in the package
structure - see `experiments/CONCEPT.MD` for the full picture:

- **Primary** (`pseudopros.Anonymizer`, below) - the full pipeline,
  including its own face detection. A transparent, drop-in replacement for
  a raw video source.
- **Secondary** (`pseudopros.Renderer`, needs `pip install pseudopros[secondary]`) -
  rendering only, no detection, for use *after* an existing CV pipeline's
  own detection/localization has already happened. See "Secondary mode"
  below.

### Common usage (Primary mode)

```python
import pseudopros

# Build the pipeline (FOMM weights auto-download ~695 MB on first call)
anon = pseudopros.Anonymizer.build(
    secret=b"your-secret-key-min-16-bytes",
    salt=b"your-deployment-salt",
)

# Open-world by default: every face process() sees is anonymized, including
# ones never explicitly enrolled - auto-enrolled on first sight so the same
# real person gets a stable synthetic identity across frames.
for frame in video_frames:
    output = anon.process(frame)  # HWC uint8 RGB, same shape as input
```

Pre-enrolling still works, and is how you find out who pseudopros is
currently tracking (auto-enrolled subjects show up here too):

```python
anon.enroll(frame_of_alice, "alice")
print(anon.enrolled_ids)  # ["alice", "auto:...", ...]
anon.unenroll("alice")
```

For a bounded set of known subjects instead - anonymize only pre-enrolled
people, leave everyone else untouched, and skip the per-face work of
deriving a style/motion state for strangers - pass `closed_roster=True`:

```python
anon = pseudopros.Anonymizer.build(
    secret=b"your-secret-key-min-16-bytes",
    salt=b"your-deployment-salt",
    closed_roster=True,
)
anon.enroll(frame_of_alice, "alice")
anon.enroll(frame_of_bob, "bob")
for frame in video_frames:
    output = anon.process(frame)  # only alice/bob get anonymized
```

Each subject keeps their own motion baseline across frames. Caution for
open-world, long-running deployments: auto-enrolled subjects accumulate in
memory for the process's lifetime (no eviction policy yet) - see
`Anonymizer`'s docstring.

### Secondary mode

For a pipeline where face detection/localization already happens upstream
(your own detector, not pseudopros's) - `Renderer` never calls a face
detector per frame; ArcFace is used once, at enrollment, for the style code:

```python
# pip install pseudopros[secondary]
from pseudopros import Renderer

renderer = Renderer.build(
    checkpoint_path="path/to/renderer_final.pt",
    secret=b"your-secret-key-min-16-bytes",
    salt=b"your-deployment-salt",
)
renderer.enroll(frame_of_alice)

# Already have a face crop from your own detector (the common case):
anonymized_crop = renderer.process(face_crop)

# Or have a full frame + bounding box instead:
anonymized_frame = renderer.process(frame, bbox=[x1, y1, x2, y2])
```

To use a local FOMM checkpoint instead of downloading:

```python
anon = pseudopros.Anonymizer.build(
    secret=b"your-secret-key-min-16-bytes",
    salt=b"your-deployment-salt",
    checkpoint_path="path/to/fomm.pth.tar",
    config_path="path/to/fomm-config.yml",
)
```

### Security validation

```python
import pseudopros

audit = pseudopros.SecurityAudit(embedder)
result = audit.audit_frame(original_frame, anonymized_frame)

print(result.is_secure)    # True if similarity < threshold
print(result.similarity)   # ArcFace cosine similarity (lower is more secure)

# Audit a sequence
summary = audit.audit_sequence(original_frames, anonymized_frames)
print(summary)
# SecurityAudit - 30 frame(s)
#   All secure:       True
#   Threshold:        0.280
#   Max similarity:   0.0031
#   Mean similarity:  -0.0018
#   Face detection:   30/30 anonymized frames

# Beyond leakage: is each subject's synthetic identity stable across frames,
# and are different subjects' synthetic identities distinguishable from
# each other? Same ArcFace-cosine-similarity methodology.
consistency = audit.audit_consistency(alice_anonymized_frames)
print(consistency.mean_similarity)  # higher = more stable pseudonym

distinctiveness = audit.audit_distinctiveness([
    alice_anonymized_frames, bob_anonymized_frames,
])
print(distinctiveness.all_distinct)  # True if every pair stayed below threshold
```

### GStreamer integration (Primary mode)

`pseudopros.gst` registers a `pseudopros` in-place video filter element -
the Primary-mode convenience wrapper, appropriate for a general-purpose
filter since it does its own detection. Requires GStreamer's Python
bindings (`python3-gi`), installed separately from pseudopros itself:

```python
import base64
import pseudopros.gst as ppgst
ppgst.register()

secret = base64.b64encode(b"your-secret-key-min-16-bytes").decode()
salt = base64.b64encode(b"your-deployment-salt").decode()

pipeline = Gst.parse_launch(
    f"v4l2src ! videoconvert ! video/x-raw,format=RGB"
    f" ! pseudopros secret={secret} salt={salt}"
    f" ! videoconvert ! autovideosink"
)
```

Add `closed-roster=true enrollment-image=/path/to/face.jpg` to the element
properties for the closed-roster/single-known-subject variant instead of
the open-world default. See `pseudopros/gst.py`'s module docstring for the
full property list.

### Configuring the synthesizer

The default synthesizer is `NoiseSynthesizer` - a fast placeholder that produces non-face-like output. Use it during development to skip the LDM model download.

For production anonymization, use `RosterSynthesizer` (downloads ~1.3 GB on first call):

```python
import pseudopros

anon = pseudopros.Anonymizer.build(
    secret=b"your-secret-key",
    salt=b"your-salt",
    synthesizer=pseudopros.RosterSynthesizer(num_steps=50),
)
```

For large known subject populations, `RapidSynthesizer` uses a pre-generated face bank for O(1) enrollment with no per-subject model inference:

```python
anon = pseudopros.Anonymizer.build(
    secret=b"your-secret-key",
    salt=b"your-salt",
    synthesizer=pseudopros.RapidSynthesizer("path/to/face_bank.npy"),
)

# Build a face bank offline (one-time, ~20s per face on CPU)
pseudopros.build_face_bank(n=10_000, output_path="face_bank.npy")
```

### Advanced / custom components

Lower-level components are importable from their submodules when you need to customize the pipeline:

```python
# Custom embedder (e.g. swap InsightFace for another ArcFace backend)
from pseudopros.embedding import InsightFaceEmbedder, FaceEmbedder, DetectedFace

embedder = InsightFaceEmbedder(model_pack="buffalo_sc")  # lighter model
anon = pseudopros.Anonymizer(model, key, embedder)

# Custom synthesizer (implement SynthesizerBackend protocol)
from pseudopros.synthesizer import SynthesizerBackend

class MyGANSynthesizer:
    def synthesize(self, seed: bytes) -> np.ndarray:
        ...  # return (256, 256, 3) float32 [0, 1]

# Access FOMM directly
from pseudopros.fomm import FommModel, FommSubjectState

model = FommModel.from_pretrained(device="cpu")            # auto-download
model = FommModel.from_checkpoint("model.pth.tar", "config.yml", device="cpu")  # local

# Per-subject stateless API — safe to interleave subjects on a shared model
state_alice = model.prepare_subject(reference_image_alice)
state_bob   = model.prepare_subject(reference_image_bob)

output_alice, state_alice = model.animate_subject(driving_frame_alice, state_alice)
output_bob,   state_bob   = model.animate_subject(driving_frame_bob,   state_bob)

# Audit result types
from pseudopros.security import AuditSummary, FrameAuditResult
```

## Command-line interface

Anonymize a video file directly from the terminal:

```bash
pseudopros anonymize input.mp4 output.mp4 \
    --enrollment face.jpg \
    --secret-file /path/to/secret.bin \
    --synthesizer roster
```

Secret options (pick one):

| Flag | Description |
|---|---|
| `--secret-file PATH` | Read raw secret bytes from a file. Recommended for production. |
| `--secret-hex HEX` | Secret as a hex string. |
| `PSEUDOPROS_SECRET` | Env var (hex). Used when no flag is provided. |

Additional options: `--salt-hex HEX`, `--device cpu\|cuda\|mps`, `--checkpoint PATH --config PATH` (to use a local FOMM checkpoint instead of downloading).

## GStreamer integration

`pseudopros.gst` provides a `pseudopros` GStreamer element that applies face anonymization as an in-place video filter. It is a separate submodule so that the rest of the library works without GStreamer installed.

### Requirements

```bash
# Debian / Ubuntu
sudo apt-get install python3-gi gstreamer1.0-python3-plugin-scanner \
    gstreamer1.0-plugins-base gstreamer1.0-plugins-good

# macOS (Homebrew)
brew install pygobject3 gstreamer
```

### Usage

Register the element once after `Gst.init()`, then use it in any pipeline string:

```python
import gi
gi.require_version("Gst", "1.0")
from gi.repository import Gst

import pseudopros.gst as ppgst

Gst.init(None)
ppgst.register()

pipeline = Gst.parse_launch(
    "v4l2src"
    " ! videoconvert ! video/x-raw,format=RGB"
    " ! pseudopros"
    "     secret=BASE64_SECRET"
    "     salt=BASE64_SALT"
    "     enrollment-image=/path/to/face.jpg"
    "     synthesizer=noise"
    " ! videoconvert ! autovideosink"
)
pipeline.set_state(Gst.State.PLAYING)
```

### Element properties

| Property | Type | Description |
|---|---|---|
| `secret` | string | Base64-encoded secret key (min 16 bytes decoded) |
| `salt` | string | Base64-encoded deployment salt (default: `pseudopros`) |
| `enrollment-image` | string | Path to an image file for subject enrollment |
| `synthesizer` | string | `noise` (fast placeholder) or `roster` (full LDM, ~1.3 GB download) |

Secrets are base64-encoded for safe embedding in pipeline strings:

```python
import base64
secret_b64 = base64.b64encode(b"your-secret-key-min-16-bytes").decode()
salt_b64 = base64.b64encode(b"your-deployment-salt").decode()
```

The element enrolls on the first frame after `enrollment-image` is set. Changing `secret` or `salt` at runtime clears the cached anonymizer and re-enrolls on the next frame.

## Running tests

```bash
# Fast tests only (no model downloads)
pytest -m "not slow"

# Full suite including security validation (~5 min)
pytest
```
