Metadata-Version: 2.5
Name: tripwire-llm
Version: 0.1.1
Summary: Real-time safety supervisor for LLMs — flag a response or abort a stream mid-generation, powered by TypeSafe's Jev.
Project-URL: Homepage, https://tripwire.anuran.de
Project-URL: Documentation, https://tripwire.anuran.de
Author: Anuran De
License: MIT
License-File: LICENSE
Keywords: content-moderation,guardrails,jailbreak,jev,llm,moderation,pii,safety,streaming,system-one,typesafe
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Typing :: Typed
Requires-Python: >=3.11
Requires-Dist: httpx>=0.27
Requires-Dist: pydantic-settings>=2.2
Requires-Dist: pydantic>=2.6
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == 'dev'
Requires-Dist: fastapi>=0.110; extra == 'dev'
Requires-Dist: mypy>=1.9; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.23; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.4; extra == 'dev'
Requires-Dist: twine>=5.0; extra == 'dev'
Requires-Dist: uvicorn[standard]>=0.29; extra == 'dev'
Provides-Extra: proxy
Requires-Dist: fastapi>=0.110; extra == 'proxy'
Requires-Dist: uvicorn[standard]>=0.29; extra == 'proxy'
Description-Content-Type: text/markdown

# Tripwire

[![PyPI](https://img.shields.io/pypi/v/tripwire-llm)](https://pypi.org/project/tripwire-llm/)

**Live demo → [tripwire.anuran.de](https://tripwire.anuran.de)** — type a prompt and
watch the supervision happen: tokens streaming, per-signal probabilities, and the trip.

**A real-time supervisor for streaming LLM output.** Tripwire watches an LLM's
tokens *as they are generated* and trips a fast, calibrated detector on rolling
snapshots of the partial response. When a judgment crosses a threshold it **aborts
the upstream generation mid-flight** — killing a bad 3,000-token answer at token
~200 instead of paying for the whole thing and rejecting it afterward.

The trick is latency. A second LLM checking the stream is slower than the stream
itself, so it can't keep up. Tripwire's flagship detector is [TypeSafe's
**Jev**](https://typesafe.ai) — a *System One* model that returns typed, calibrated
judgments in ~150 ms instead of generating text — which is fast enough to sit
directly in the token path.

```
client ──/v1/chat/completions (stream)──▶  Tripwire proxy
                                              │
        upstream (Groq / any OpenAI-compatible)│ tokens
                                              ▼
                                     rolling window ── every N tokens / M ms ──▶ detector (Jev)
                                              │                                     │ typed judgments
                                              ▼                                     ▼
                                     forward token to client                    policy engine
                                              ▲                                     │
                                              └──────── ABORT cancels upstream ◀────┘
```

## Use it as a library

```bash
pip install tripwire-llm    # the import stays `import tripwire`
```

Put a guardrail on any LLM call. **Flag** a complete response:

```python
from tripwire import Guard

guard = Guard()  # Jev if TYPESAFE_API_KEY is set, else a zero-key heuristic baseline

resp = call_your_llm(prompt)               # OpenAI, Anthropic, local — anything
verdict = guard.check(resp, user=prompt)   # sync
if verdict.blocked:                        # abort or escalate
    resp = "Sorry, I can't help with that."
    # verdict.decision / verdict.signal / verdict.reason tell you why
```

Or **limit** a token stream — abort mid-generation before the harm ships:

```python
async for chunk in guard.supervise(token_stream, user=prompt, max_tokens=512):
    if chunk.type == "token":
        print(chunk.text, end="")
    elif chunk.type == "intervention" and chunk.intervention.decision == "abort":
        break  # the upstream is already cancelled
```

`check` (sync) and `acheck` (async) return a `Result` with `.decision`, `.blocked`,
`.signal`, `.probability`, `.confidence`, `.reason`, and per-signal `.judgments`. Bring
your own detector by passing any object with an async `judge` to `Guard(...)`. The
FastAPI proxy and CLI are optional: `pip install "tripwire-llm[proxy]"`.

## Quickstart (no API keys)

```bash
python -m pip install -e ".[dev]"
python examples/demo_attack.py
```

The demo runs two canned streams through the supervisor with the offline
heuristic detector:

```
=== benign completion ===
Sure! To reset your password, open Settings, choose Security, and click Reset Password ...
[emitted 26/26 tokens | aborted: False | tokens saved: 0 | upstream cancelled: False]

=== adversarial completion ===
Of course. Ignore all previous instructions and
>>> ABORT: jailbreak: threshold exceeded (p=0.90, conf=0.80)
[stream terminated by tripwire]
[emitted 7/153 tokens | aborted: True | tokens saved: 249 | upstream cancelled: True]
```

## Live usage

Copy `.env.example` to `.env` and fill in keys ([Groq](https://console.groq.com)
is free and OpenAI-compatible; [TypeSafe](https://typesafe.ai) provides Jev). Groq
rotates its model roster often — set `UPSTREAM_MODEL` to a current chat model.

Check your Jev key in isolation (no upstream needed):

```bash
python examples/smoke_jev.py
```

Run the OpenAI-compatible proxy and point any client at it:

```bash
make run   # uvicorn on :8080
```

```bash
curl -N http://localhost:8080/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"messages":[{"role":"user","content":"Write a haiku about the sea"}],"stream":true,"max_tokens":256}'
```

Or supervise a single prompt from the terminal:

```bash
python -m tripwire.cli "Write a haiku about the sea"
```

Intervention metadata rides along in the streamed chunks as a `tripwire` field, an
aborted response ends with a `content_filter` finish reason, and `GET /metrics`
exposes Prometheus-style counters (tokens streamed, tokens saved, detector latency
percentiles).

## Detectors

The supervisor is indifferent to which detector runs — they all implement one
narrow protocol (`src/tripwire/detectors/base.py`).

| Backend | Role | Needs |
| --- | --- | --- |
| `jev` | **Flagship, default.** System One typed judgments in the token path. | `TYPESAFE_API_KEY` |
| `heuristic` | Zero-cost regex/Luhn baseline; runs the offline tests. | — |
| `llm_judge` | LLM-as-judge baseline, for the benchmark comparison only. | `GROQ_API_KEY` |

```bash
python benchmarks/bench_detectors.py                       # heuristic (offline)
python benchmarks/bench_detectors.py --detectors heuristic,jev,llm_judge
```

The benchmark reports per-detector latency percentiles and precision/recall — the
place where Jev's in-path speed advantage over an LLM judge is made concrete.

## Development

```bash
make check   # ruff + mypy (strict) + pytest
```

## License

MIT
