Metadata-Version: 2.5
Name: scrutineer-agents
Version: 0.3.0
Summary: Agent Behavioral Testing Platform — tests what agents DO, not just what they SAY
Project-URL: Homepage, https://github.com/adam85sims/scrutineer
Project-URL: Documentation, https://github.com/adam85sims/scrutineer#readme
Project-URL: Repository, https://github.com/adam85sims/scrutineer
Project-URL: Issues, https://github.com/adam85sims/scrutineer/issues
Author: Adam Sims
License-Expression: MIT
License-File: LICENSE
Keywords: agent,ai,behavioral,chaos,evaluation,testing
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.11
Requires-Dist: click>=8.0
Provides-Extra: adapters
Requires-Dist: crewai>=0.10.0; extra == 'adapters'
Requires-Dist: langchain-core>=0.3.0; extra == 'adapters'
Requires-Dist: openai>=1.0.0; extra == 'adapters'
Provides-Extra: all
Requires-Dist: crewai>=0.10.0; extra == 'all'
Requires-Dist: fastapi>=0.115.0; extra == 'all'
Requires-Dist: httpx>=0.27.0; extra == 'all'
Requires-Dist: langchain-core>=0.3.0; extra == 'all'
Requires-Dist: openai>=1.0.0; extra == 'all'
Requires-Dist: opentelemetry-exporter-otlp>=1.0.0; extra == 'all'
Requires-Dist: opentelemetry-sdk>=1.0.0; extra == 'all'
Requires-Dist: pytest-asyncio; extra == 'all'
Requires-Dist: pytest-cov>=4.0; extra == 'all'
Requires-Dist: pytest-playwright>=0.5.0; extra == 'all'
Requires-Dist: pytest-randomly>=4.0; extra == 'all'
Requires-Dist: pytest-xdist>=3.0; extra == 'all'
Requires-Dist: pytest>=7.0; extra == 'all'
Requires-Dist: pyyaml>=6.0; extra == 'all'
Requires-Dist: ruff>=0.1.0; extra == 'all'
Requires-Dist: sse-starlette>=2.0.0; extra == 'all'
Requires-Dist: uvicorn[standard]>=0.34.0; extra == 'all'
Provides-Extra: crewai
Requires-Dist: crewai>=0.10.0; extra == 'crewai'
Provides-Extra: dev
Requires-Dist: httpx>=0.27.0; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.23.0; extra == 'dev'
Requires-Dist: pytest-cov>=4.0; extra == 'dev'
Requires-Dist: pytest-playwright>=0.5.0; extra == 'dev'
Requires-Dist: pytest-randomly>=4.0; extra == 'dev'
Requires-Dist: pytest-xdist>=3.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff>=0.1.0; extra == 'dev'
Provides-Extra: governance
Requires-Dist: pyyaml>=6.0; extra == 'governance'
Provides-Extra: langchain
Requires-Dist: langchain-core>=0.3.0; extra == 'langchain'
Provides-Extra: openai
Requires-Dist: openai>=1.0.0; extra == 'openai'
Provides-Extra: otel
Requires-Dist: opentelemetry-exporter-otlp>=1.0.0; extra == 'otel'
Requires-Dist: opentelemetry-sdk>=1.0.0; extra == 'otel'
Provides-Extra: web
Requires-Dist: fastapi>=0.115.0; extra == 'web'
Requires-Dist: httpx>=0.27.0; extra == 'web'
Requires-Dist: sse-starlette>=2.0.0; extra == 'web'
Requires-Dist: uvicorn[standard]>=0.34.0; extra == 'web'
Description-Content-Type: text/markdown

# Scrutineer

**Agent Behavioral Testing Platform** — Tests what agents DO, not just what they SAY.

[![Python 3.11+](https://img.shields.io/badge/python-3.11+-blue.svg)](https://www.python.org/downloads/)
[![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
[![Tests](https://img.shields.io/badge/tests-520%20passing-brightgreen.svg)](#testing)
[![Code style: ruff](https://img.shields.io/badge/code%20style-ruff-red.svg)](https://docs.astral.sh/ruff/)

## The Problem

88% of AI agents fail in production. The dominant failure modes are operational:
- Tool errors (28%)
- Memory/state issues (22%)
- Edge cases (18%)

Yet the entire evaluation ecosystem (DeepEval, LangSmith, MS AGT) focuses on
output quality or observability. Nobody tests agent **behavior** in production-like
environments before deployment.

## The Solution

Scrutineer fills that gap. It's a behavioral testing platform that:

1. **Mocks your agent's environment** — tools, APIs, databases with configurable
   latency, errors, and rate limits
2. **Injects chaos** — tool failures, context degradation, cascading errors,
   spec drift under pressure
3. **Asserts behavior** — 20+ assertions across tool calls, state consistency,
   governance compliance, resilience, and performance
4. **Reports regressions** — structural diffing, baseline comparison, HTML + JUnit reports

## Quick Start

> **Distribution name:** this project publishes to PyPI as **`scrutineer-agents`**.
> Do **not** run `pip install scrutineer` — that name belongs to an unrelated project and
> installs a different library. The import package and the console script are both
> `scrutineer`.

```bash
# Install (zero dependencies by default)
pip install scrutineer-agents

# Or with framework adapters
pip install "scrutineer-agents[adapters]"

# Or with the WebUI dashboard
pip install "scrutineer-agents[web]"

# Run the quickstart example
python examples/langchain_quickstart.py

# Run a YAML scenario
scrutineer run --path examples/basic_scenario.yaml
```

## WebUI Dashboard

Scrutineer includes a browser-based dashboard for running scenarios, viewing
traces, and comparing baselines — all wrapping the core Python API.

```bash
# Install with web dependencies
pip install "scrutineer-agents[web]"

# Start the dashboard
scrutineer serve

# Or with a custom port
scrutineer serve --port 9090
```

Then open [http://localhost:8080](http://localhost:8080) in your browser.

Features:
- **Dashboard** — pass/fail stats, recent runs, quick actions
- **Scenarios** — browse, inspect, and run test scenarios
- **Runs** — live execution with step-by-step trace visualization
- **Baselines** — saved results with regression diff comparison
- **Live Console** — real-time log streaming via SSE during test runs

See [src/scrutineer/web/README.md](src/scrutineer/web/README.md) for the full
WebUI guide.

## Architecture

```
src/scrutineer/
├── env.py          # MockTool, MockAPI, MockDatabase, EnvironmentBuilder
├── chaos.py        # ToolFailureInjector, ContextDegradation, CascadingFailures
├── assertions.py   # 20+ behavioral assertions
├── runner.py       # @scrutineer_test decorator, ScenarioRunner
├── reporting.py    # Regression reports, JUnit XML, HTML
├── baseline.py     # JSON baseline storage with git integration
├── otel.py         # OpenTelemetry span model
├── cli.py          # Full CLI: run, list, info, baseline, diff, report
├── adapters/       # LangChain, CrewAI, OpenAI SDK, Generic
└── web/            # FastAPI WebUI dashboard (optional)
    ├── app.py          # FastAPI application factory
    ├── server.py       # Uvicorn entry point
    ├── api/            # REST API routers (scenarios, runs, baselines)
    ├── services/       # Service layer wrapping core modules
    ├── schemas/        # Pydantic request/response models
    └── static/         # Frontend (HTML, CSS, JS)
```

## The Chaos Module (Differentiator)

Scrutineer's chaos injection is what sets it apart:

- **ContextDegradation** — Quadratic acceleration curve matching real context
  window pressure (last 20% is much worse than first 20%)
- **CascadingFailures** — Multi-agent error propagation with dependency graphs
  (database → api_server → ui)
- **SpecDrift** — Agent improvisation under pressure with intensity levels
  and cumulative drift scoring

No other tool tests these production failure modes.

## CLI Commands

```bash
scrutineer run <scenario>          # Run a test scenario
scrutineer list                    # List available scenarios
scrutineer info <scenario>         # Show scenario details
scrutineer baseline record         # Record current state as baseline
scrutineer baseline show           # Show recorded baseline
scrutineer diff                    # Compare current vs baseline
scrutineer report                  # Generate regression report
scrutineer trace <run-id>          # Show execution trace
scrutineer serve                   # Start WebUI dashboard
scrutineer serve --port 9090       # Custom port
```

## Framework Adapters

Scrutineer ships adapters for specific frameworks, plus a generic hook adapter for anything else:

```python
# LangChain — rebinds your_agent.tools so the agent's own call path hits the mocks
from scrutineer.adapters.langchain import wrap_agent
wrapped = wrap_agent(your_agent, tool_map={...}, trace=trace)

# Agents whose tools are bound internally (a create_react_agent Runnable, say)
# cannot be rebound. wrap_agent raises AgentInterceptionError rather than
# quietly letting the real tools run — build the agent against the mocks instead:
wrapped = wrap_agent(agent=None, tool_map={...}, trace=trace, intercept=False)
agent = create_react_agent(model, wrapped.tools.values())

# CrewAI
from scrutineer.adapters.crewai import wrap_crew_agent
wrapped = wrap_crew_agent(your_crew, tool_map={...}, trace=trace)

# OpenAI SDK
from scrutineer.adapters.openai import wrap_agent
wrapped = wrap_agent(your_agent, tool_map={...}, trace=trace)

# Generic (any framework)
from scrutineer.adapters.generic import HookAdapter
adapter = HookAdapter(mock=your_mock, before=hook_fn)
```

## Chaos Example

```python
from scrutineer.chaos import (
    ToolFailureInjector,
    ContextDegradation,
    CascadingFailures,
    ChaosBudget,
)

# Fail 30% of tool calls with timeout errors
injector = ToolFailureInjector(
    failure_type="timeout",
    probability=0.3,
)

# Degrade context with quadratic acceleration
degradation = ContextDegradation(strategy="TRUNCATION")

# Cascade failures from database to API to UI using a custom dependency graph
cascade = CascadingFailures(
    cascade_probability=0.7,
    max_cascade_depth=3,
    dependency_graph={
        "database": "api_server",
        "api_server": "user_interface",
    },
)

# Cap total failures per run
budget = ChaosBudget(max_failures=10)
```

## Documentation

- [Quickstart](docs/QUICKSTART.md) — 5-minute guide from install to first test
- [Chaos Guide](docs/CHAOS_GUIDE.md) — Deep dive on failure injection patterns
- [Adapters Guide](docs/ADAPTERS_GUIDE.md) — How to write custom adapters
- [Integration Testing](docs/INTEGRATION_TESTING.md) — Proof of value with real LangChain tools
- [API Reference](docs/api.md) — Module documentation
- [WebUI Design](docs/WEBUI_DESIGN.md) — Architecture and implementation plan
- [WebUI Guide](src/scrutineer/web/README.md) — Getting started with the dashboard

## Testing

```bash
# Run all tests
pytest tests/ -v

# Run with coverage
pytest tests/ --cov=scrutineer --cov-report=html

# Run integration tests only
pytest tests/scrutineer/test_integration_langchain.py -v

# Lint
ruff check src/ tests/
```

## License

MIT
