Metadata-Version: 2.4
Name: antipluto
Version: 1.0.0
Summary: A unified framework for email preprocessing, PII masking, LLM cohort generation, and dual-branch semantic-stylometric phishing classification.
License: CC BY-NC 4.0
Classifier: Programming Language :: Python :: 3
Classifier: License :: Free for non-commercial use
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Security
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: click>=8.0
Requires-Dist: pydantic>=2.0
Requires-Dist: PyYAML>=6.0
Requires-Dist: beautifulsoup4>=4.12
Requires-Dist: spacy>=3.7
Requires-Dist: langdetect>=1.0.9
Requires-Dist: openai>=1.0
Requires-Dist: anthropic>=0.30
Requires-Dist: pandas>=2.0
Provides-Extra: classifier
Requires-Dist: scikit-learn>=1.4; extra == "classifier"
Requires-Dist: scipy>=1.12; extra == "classifier"
Requires-Dist: xgboost>=2.0; extra == "classifier"
Requires-Dist: accelerate>=0.30; extra == "classifier"
Requires-Dist: datasets>=2.0; extra == "classifier"
Requires-Dist: sentencepiece>=0.2.0; extra == "classifier"
Requires-Dist: tiktoken>=0.7.0; extra == "classifier"
Requires-Dist: onnx>=1.16; extra == "classifier"
Requires-Dist: onnxruntime>=1.18; extra == "classifier"
Requires-Dist: onnxscript>=0.1.0; extra == "classifier"
Requires-Dist: skl2onnx>=1.17; extra == "classifier"
Requires-Dist: onnxmltools>=1.12; extra == "classifier"
Requires-Dist: optimum>=1.20; extra == "classifier"
Requires-Dist: matplotlib>=3.8; extra == "classifier"
Provides-Extra: api
Requires-Dist: fastapi>=0.110.0; extra == "api"
Requires-Dist: uvicorn[standard]>=0.28.0; extra == "api"
Requires-Dist: python-multipart>=0.0.9; extra == "api"
Provides-Extra: dev
Requires-Dist: dvc[gdrive]; extra == "dev"
Requires-Dist: spacy-transformers>=1.3; extra == "dev"
Dynamic: license-file

<div align="center">
  <img src="antipluto.svg" alt="Anti-PLUTO Logo" width="120" />
  <h1>Anti-PLUTO</h1>
  <p><strong>Anti-Phishing Lexical Utilities and Threat Observation</strong></p>
  <p>
    <em>A unified framework for email dataset preprocessing, PII masking, synthetic data generation, and semantic-stylometric phishing classification.</em>
  </p>
  <p>
    <a href="https://github.com/mmaarij/antipluto/blob/main/LICENSE">
      <img src="https://img.shields.io/badge/License-CC%20BY--NC%204.0-mistyrose.svg" alt="License: CC BY-NC 4.0" />
    </a>
    <a href="https://python.org">
      <img src="https://img.shields.io/badge/Python-3.10%2B-lightblue" alt="Python 3.10+" />
    </a>
  </p>
</div>

<hr />

## Overview

With the rapid proliferation of Large Language Models (LLMs) enabling threat actors to generate contextually fluent spear-phishing emails that evade traditional rule-based filters (like SpamAssassin and Rspamd), defensive systems must adapt. **Anti-PLUTO** addresses this emerging threat by utilizing a state-of-the-art dual-branch feature fusion architecture.

## Key Components & Architecture

### 1. Dual-Branch Feature Fusion
- **High-Dimensional Stylometric Branch (100,000 dimensions)**: Extracts joint word- and character-level $n$-gram TF-IDF representations. It preserves functional stop-words, capitalization patterns, and punctuation marks to isolate the subtle, subconscious stylistic fingerprints of human writers vs. LLM token sampling distributions.
- **Deep Contextual Semantic Branch (768 dimensions)**: Utilizes a dense continuous vector manifold generated by fine-tuning DeBERTa-v3-small (Decoding-enhanced BERT with Disentangled Attention). It extracts contextual intent, emotional manipulation vectors, and social engineering semantics.

### 2. Unified 4-Cohort Classification
Anti-PLUTO is designed to effectively differentiate between four email cohorts:
- **Human Benign (HB)**: Legitimate, human-authored corporate correspondence.
- **Human Phishing (HP)**: Malicious, human-authored attacks (e.g., credential harvesting, advance-fee fraud).
- **LLM Benign (LB)**: Legitimate corporate communications written or polished by generative writing assistants.
- **LLM Phishing (LP)**: Malicious communications synthesized or polished by generative AI to execute social engineering.

### 3. Privacy-Preserving Lexical Processing Pipeline
Features an extensive 5-stage preprocessing pipeline designed to prevent target leakage:
- **RFC 5322 MIME Extraction**: Parses complex multipart email structures.
- **HTML Stripping & Thread Slicing**: Cleans markup and removes historic reply chains.
- **Unicode Normalization**: Mitigates homoglyph attacks.
- **Two-Stage PII Masking**: Employs deterministic regex tokenization and spaCy NER to mask URLs, IPs, email addresses, names, and organizations into standard tokens (e.g., `[NAME]`, `[URL]`).

### 4. Production Serving
Includes an asynchronous FastAPI REST server supporting inference on pre-computed feature vectors at high throughput (up to 7,497 msgs/s), with ONNX model export support.

---

## Installation

Anti-PLUTO is built with modularity in mind. You can install the core framework or include optional components like the classifier and REST API depending on your needs.

### Quick Setup (Recommended)
For the most streamlined installation experience, we have provided a setup script that installs the core framework, all optional dependencies, and the required language models in an editable state.

```bash
git clone https://github.com/mmaarij/antipluto.git
cd antipluto
chmod +x setup_env.sh
./setup_env.sh
```

### Standard Installation

Once published to PyPI, you can install the core framework directly:

```bash
pip install antipluto
```

To install directly from the source repository:

```bash
git clone https://github.com/mmaarij/antipluto.git
cd antipluto
pip install -e .
```

#### Optional Dependencies

You can install specific components of the framework depending on your use case:

- **Classifier Mode:** `pip install -e ".[classifier]"` (Installs ML dependencies)
- **API Mode:** `pip install -e ".[api]"` (Installs FastAPI and Uvicorn for REST endpoints)
- **Development/Full:** `pip install -e ".[dev,classifier,api]"`

#### Language Models (Required for Masking)

Because direct URL dependencies are restricted on PyPI, you must install the required spaCy language models manually after installing the framework:

```bash
# Standard model (Fast, recommended for general usage)
pip install https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.8.0/en_core_web_sm-3.8.0-py3-none-any.whl

# Transformer model (Highest accuracy, requires 'dev' extra dependencies)
pip install https://github.com/explosion/spacy-models/releases/download/en_core_web_trf-3.8.0/en_core_web_trf-3.8.0-py3-none-any.whl
```

### GPU Acceleration (Optional)

To enable GPU acceleration for the transformer-based masking and classification, you must install PyTorch with CUDA support manually. Ensure you have the CUDA Toolkit 13.x installed, then run:

```bash
pip install cupy-cuda13x
pip install torch --index-url https://download.pytorch.org/whl/cu132 --force-reinstall
```

---

## Configuration

Before running the framework, you must configure your API keys (if you are generating datasets) and other system settings. 

1. Copy the example configuration file:
   ```bash
   cp antipluto/configs/default.example.yaml antipluto/configs/default.yaml
   ```
2. Open `antipluto/configs/default.yaml` and fill in your respective keys (e.g., `OPENROUTER_API_KEY`, `FOUNDRY_API_KEY`).
   
*Note: The `default.yaml` file is intentionally ignored by git to prevent sensitive data leakage.*

---

## Usage

Anti-PLUTO provides a robust CLI interface. You can invoke it after installation as `antipluto` or via python as `python -m antipluto`.

### 1. Preprocessing Data
Run the full ingestion pipeline to clean and process your raw corpora:
```bash
# Run with default configuration:
antipluto preprocess

# Run a specific dataset source (e.g., nazario):
antipluto preprocess --source nazario

# Override the export mode (full = 20 features, minimal = core NLP):
antipluto preprocess --mode full
```

### 2. PII Masking
Apply Named Entity Recognition (NER) and regex masking to anonymize sensitive information and prevent target leakage.

```bash
# Mask all preprocessed cohorts:
antipluto mask --input datasets/datasets_preprocessed/

# Mask a specific file:
antipluto mask --input datasets/datasets_preprocessed/human_written_phishing.jsonl

# High-accuracy transformer masking (requires GPU and en_core_web_trf):
antipluto mask --input datasets/datasets_preprocessed/ --spacy-model en_core_web_trf --batch-size 64
```

### 3. LLM Cohort Generation
Generate synthetic LLM cohorts using Azure OpenAI, Anthropic, or OpenRouter models (configurable via `default.yaml`).
```bash
# Generate benign and phishing cohorts:
antipluto generate --cohort benign
antipluto generate --cohort phishing

# Trace a synthetic email back to its original human seed email:
antipluto trace --record '{"source": "gpt-5-mini", "seed_idx": 42, "label": 1}'
```

### 4. Training the Classifiers
Train the dual-branch feature fusion classifier with a chosen fusion head (XGBoost, Random Forest, PyTorch MLP, 1D-CNN).
```bash
# Train the default XGBoost model on all cohorts:
antipluto train --classifier xgboost

# Compare all available models (XGBoost, RF, MLP, CNN) to find the best configuration:
antipluto compare --classifier all
```

### 5. Prediction & Exporting
Classify individual emails and export trained models to ONNX for portable deployment.
```bash
# Predict a single email's cohort and phishing probability:
antipluto predict --subject "Urgent Account Update" --body "Click here to secure your account."

# Predict from a raw email file using a specific fusion head:
antipluto predict --file suspicious_email.eml --classifier rf --mode fusion

# Export trained DeBERTa and classifier heads to ONNX format:
antipluto export --classifier all
```

### 6. Baseline Benchmarking
Evaluate Anti-PLUTO against industry-standard legacy engines (SpamAssassin & Rspamd) on unseen test emails.
```bash
antipluto baseline --engine all
```

### 7. Production API Serving
Launch the asynchronous FastAPI inference server to expose REST endpoints for real-time email security prediction.
```bash
# Start the API server on 127.0.0.1:8000
antipluto serve --classifier xgboost --mode fusion

# Enable auto-reload for local development:
antipluto serve --reload
```

### 8. Validation & Telemetry
Evaluate the statistics and schema integrity of your generated cohorts.
```bash
antipluto validate --input datasets/datasets_preprocessed/human_written_benign.jsonl
antipluto stats --input datasets/datasets_preprocessed/human_written_phishing.jsonl
```
