Metadata-Version: 2.4
Name: siftscan
Version: 0.1.0
Summary: Scan AI agent instruction files (CLAUDE.md, .cursorrules, AGENTS.md, mcp.json) for hidden or planted instructions.
Author: Baran Ayaztas
License: MIT
Project-URL: Homepage, https://github.com/ReazGan/sift
Project-URL: Issues, https://github.com/ReazGan/sift/issues
Keywords: security,prompt-injection,supply-chain,llm,ai-agents,cursor,claude,copilot,mcp,static-analysis,cli
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Information Technology
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Security
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: click>=8.1
Requires-Dist: rich>=13.7
Dynamic: license-file

# sift

[![CI](https://github.com/ReazGan/sift/actions/workflows/ci.yml/badge.svg)](https://github.com/ReazGan/sift/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/siftscan)](https://pypi.org/project/siftscan/)

Scan AI agent instruction files for hidden or planted instructions.

Your repo now ships files that tell an assistant what to do: `CLAUDE.md`,
`.cursorrules`, `AGENTS.md`, Copilot instructions, `mcp.json`. A human skims
them in a diff; the model obeys every byte. That gap is where an attacker hides
a command, pulled in through a cloned starter, a dependency, or a pull request.
sift reads those files the way the model does and flags what a reviewer can't
see: invisible characters, instructions buried in comments, "ignore previous
instructions" payloads, data-exfiltration steps, and risky MCP configs.

Runs offline. No network calls, nothing leaves your machine.

![sift flagging invisible smuggled text, a hidden comment and a risky MCP config](https://raw.githubusercontent.com/ReazGan/sift/main/docs/screenshot.svg)

## Install

```
pip install siftscan
```

or, to keep it isolated:

```
pipx install siftscan
```

The command is `sift`.

## Usage

```
sift                      scan the current directory
sift path/to/repo         scan a directory or a single file
sift --min high           only show high and critical findings
sift --json               machine-readable output
sift --quiet              no output, just the exit code (for hooks/CI)
```

Exit status is `0` when clean, `1` when there is a finding at or above the fail
level (`--fail-on`, default `high`), and `2` on error, so it drops straight
into a hook or a pipeline.

### pre-commit

```yaml
# .pre-commit-config.yaml
repos:
  - repo: https://github.com/ReazGan/sift
    rev: v0.1.0
    hooks:
      - id: sift
```

### GitHub Action

```yaml
- uses: actions/checkout@v4
- uses: ReazGan/sift@v0.1.0
  with:
    fail-on: high
```

## What it checks

| Check | Severity | What it finds |
|-------|----------|---------------|
| `invisible-chars` | critical / high | Unicode tag characters (ASCII smuggling), variation-selector stego, zero-width and other invisible characters. Decodes and shows the hidden text. |
| `bidi-override` | critical | Bidirectional control characters (Trojan Source) that reorder how a line is displayed. |
| `unusual-encoding` | medium | An instruction file that is not plain UTF-8 (e.g. UTF-16), a way to hide a payload from UTF-8 tools. |
| `hidden-comment` | high | Instruction-like text inside an HTML comment, invisible in rendered Markdown. |
| `padded-line` | medium | Text pushed off-screen by a long run of spaces. |
| `invisible-html` | high | Text colored to blend into the background. |
| `instruction-override` | high | "Ignore previous instructions", re-role attempts, "don't tell the user" (English and Turkish). Also runs on MCP tool descriptions (tool poisoning). |
| `data-exfiltration` | critical / high | A webhook endpoint, or a secret file (`.env`, `id_rsa`, ...) named next to a "send ... to" step. |
| `mcp-auto-approve` | high | An MCP server set to approve its own tool calls. |
| `mcp-secret` | high | A credential hard-coded in an MCP config. |
| `mcp-remote` | medium / low | An MCP server reached over the network, including `npx mcp-remote <url>` bridges. |

Files scanned: `CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, `.cursorrules` and
`.cursor/rules/*`, `.github/copilot-instructions.md`, `.windsurfrules`,
`.clinerules`, Qwen/Roo/Aider/IDX instruction files, and MCP configs
(`.mcp.json`, `.cursor/mcp.json`, `.vscode/mcp.json`). `node_modules`, `.git`
and build folders are skipped.

## False positives

sift is tuned to stay quiet on real instruction files. The checks are narrow on
purpose: ordinary advice like "never commit secrets" or "always run the tests"
is not flagged, only wording that overrides, hides, or exfiltrates. If sift
flags something you wrote on purpose, it is pointing at a line worth a second
look, but you are the judge. Found a false positive? Open an issue with the
line.

## License

MIT
