Metadata-Version: 2.4
Name: internal-affairs-4-ai
Version: 0.1.0
Summary: AI agent forensics & cost governance. Instrument agents, reconstruct why they chose what they did, and audit whether a cheaper model could do the same job.
Author: Internal Affairs 4 AI
License-Expression: Apache-2.0
Keywords: llm,agents,observability,forensics,cost-optimization,langgraph
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Debuggers
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pydantic<3,>=2.7
Requires-Dist: pydantic-settings<3,>=2.2
Requires-Dist: fastapi<1,>=0.110
Requires-Dist: uvicorn[standard]<1,>=0.29
Provides-Extra: langgraph
Requires-Dist: langgraph>=0.2; extra == "langgraph"
Provides-Extra: opentelemetry
Requires-Dist: opentelemetry-api>=1.25; extra == "opentelemetry"
Requires-Dist: opentelemetry-sdk>=1.25; extra == "opentelemetry"
Requires-Dist: opentelemetry-proto>=1.25; extra == "opentelemetry"
Provides-Extra: oidc
Requires-Dist: cryptography>=42; extra == "oidc"
Provides-Extra: postgres
Requires-Dist: psycopg[binary]>=3.1; extra == "postgres"
Provides-Extra: presidio
Requires-Dist: presidio-analyzer>=2.2; extra == "presidio"
Provides-Extra: mcp
Requires-Dist: mcp<2,>=1.2; extra == "mcp"
Provides-Extra: ui
Requires-Dist: playwright>=1.40; extra == "ui"
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: pytest-cov>=5.0; extra == "dev"
Requires-Dist: httpx>=0.27; extra == "dev"
Requires-Dist: ruff>=0.4; extra == "dev"
Requires-Dist: cryptography>=42; extra == "dev"
Dynamic: license-file

# Internal Affairs 4 AI 🕵️

AI agent forensics & cost governance. Instrument agents, reconstruct **why** they
chose what they did — and **cut cost by swapping expensive models for cheaper ones
that produce the same result**, per job, per tool, per action.

> "Police investigation" for autonomous systems: behavioral forensics + a budget audit.

## What it does

1. **Connect to any agent system** (LangGraph first, then any OpenTelemetry GenAI source
   via the OTLP endpoint).
2. **Reconstruct the case file** — what the job was, what the agent *claimed* to do,
   what it *actually* did, and what independent artifacts prove it
   (the **claim / behavior / provenance** triad).
3. **Emit findings** with an explicit epistemic level, so we never present a
   hypothesis as an observation.
4. **Audit cost** and recommend when a cheaper model is a credible candidate — and,
   just as importantly, say **"no clear winner"** when the evidence is weak.
5. **Diff two runs** (git-for-agents) to see exactly which decisions, tools, and
   models changed — and classify the **root cause** of a failed or anomalous run.
6. **Enforce policies** — configurable governance rules (max cost, PII presence,
   expensive-model justification, tool allow/deny lists) that emit violations as findings.
7. **Evaluate rigorously** — multi-seed golden-set evaluation (≥3 seeds) with mean ± std,
   so the "cheaper model?" verdict is statistically honest, not a single-run guess.
8. **Evaluate live** — plug a real model client (OpenAI-compatible) into the evaluator and
   answer "is gpt-4o-mini good enough for this job?" against live models.
9. **Stay current** — a syncable model catalog (`ia catalog --sync <url>`) keeps prices,
   tiers, context windows, and new models up to date.
10. **Authenticate & authorize** — JWT + API-key + **OIDC (RS256/JWKS)** auth with RBAC roles
    (`admin` / `investigator` / `viewer`).
11. **Route intelligently** — cascade recommender: try the cheap model first, escalate to the
    expensive one on low confidence, with projected savings.
12. **Judge safely** — an injection-screened, structured-input LLM judge (sandboxed, no tools)
    that treats its own output as *evidence*, never a verdict.
13. **Ingest asynchronously** — OTLP batches are store-and-ack'd; a background worker builds
    case files off the request path and enforces retention (`IA_TRACE_RETENTION`).
14. **Reconcile the bill** — compare metered cost against Anthropic/OpenAI usage reports to
    surface unaccounted (shadow) usage.
15. **Alert on findings** — HMAC-signed webhooks fire when a case meets a severity threshold.

## The core value: per-job, tool-aware cost cutting

For every job, the system reads **which tools the agent used** and **which models it
called**, then recommends the cheapest model that can produce the *same result*:

- **Capability-matched** — filters candidates by tool-calling support, context window,
  and deprecation, so you never swap to a model that can't do the job.
- **Token-mix priced** — candidates are ranked by projected cost on the job's *actual*
  input/output token counts, not list prices: an input-heavy job (RAG, long-document
  summarization) ranks input price first; a generation-heavy job ranks output price first.
- **Quality-verified** — with a golden set + runner, each candidate is run through the
  multi-seed evaluator; a cheaper model that *fails the task* is skipped, and the
  cheapest tied-or-better candidate wins.
- **Free included** — zero-cost models (e.g. NVIDIA NIM open models) are ranked first,
  so "use the free model" is the default suggestion when it holds up.

```python
from internal_affairs.verdict.cost_cutter import analyze_cost_cuts

report = analyze_cost_cuts(case)  # capability-matched
report = analyze_cost_cuts(case, golden_set=golden, runner=runner, seeds=5)  # quality-verified
```

`GET /case/{id}/cost-cuts` returns the same report for the dashboard.

## Security-first

Redaction happens **at ingest, before storage**: secrets and PII never reach the
store or the (future) LLM judges. A tamper-evident, hash-chained, signed evidence
log detects any post-hoc alteration. See [`SECURITY.md`](SECURITY.md) and
[`docs/threat-model.md`](docs/threat-model.md).

## Install the server (Python)

The server is a plain Python package. It installs a console script also named
`ia` (`ia serve`, `ia demo`, `ia catalog`) — the Go client is a *different*
program with the same name, covered in [CLI (Go client)](#cli-go-client) below.

```bash
uv venv && source .venv/bin/activate
uv pip install -e ".[dev]"          # core + dev/test deps
uv pip install -e ".[langgraph]"    # optional: LangGraph adapter
```

## Run the demo

```bash
python -m examples.demo     # full pipeline: record → redact → investigate → verdict
ia catalog                 # list the model catalog
ia catalog --sync <url>    # merge new models/prices from a remote registry
```

## Run the API

```bash
cp .env.example .env        # set IA_EVIDENCE_SIGNING_KEY in prod
ia serve                    # or: uvicorn internal_affairs.api.app:app --reload
```

**Web UI:** open `http://127.0.0.1:8000/` for a dashboard (cases, findings/verdicts,
model catalog, audit log, activity feed, run graph, terminal) served by FastAPI — no
build step.

## CLI (Go client)

A Go client for the same API — `ia` — in [`cli/`](cli/README.md). It is a
client, never a server and never a shell: every command maps onto one REST route,
and the server decides what a credential may do.

> **`ia` is two different programs.** The Python package installs a console
> script called `ia` (`ia serve`, `ia demo`, `ia catalog` …) — that is the
> *server*. This Go binary is also called `ia` (`ia doctor`, `ia case`, `ia run`,
> `ia smoke-test` …) — that is the *client*. Whichever directory comes first on
> `$PATH` wins, so pick one to own the name, or rename one:
> `make -C cli build && mv cli/bin/ia /usr/local/bin/iactl`.

**Option 1 — install a released binary (no Go toolchain needed).** Recommended:

```bash
curl -fsSL https://raw.githubusercontent.com/smoolzone/ia4ai/main/cli/scripts/install.sh | sh
```

The script detects `linux`/`darwin` × `amd64`/`arm64`, installs to `/usr/local/bin`
(or `~/.local/bin` when that is not writable), and **refuses to install unless the
archive's SHA256 matches the release's `SHA256SUMS`**.

> **This repository is private**, and GitHub answers `404` (not `403`) to a client
> without a token — for the raw script, the release assets and the API alike. Export
> a token with read access first and pass it on the script fetch:

```bash
export GITHUB_TOKEN=...   # any token with read access to this repo
curl -fsSL -H "Authorization: Bearer $GITHUB_TOKEN" \
  https://raw.githubusercontent.com/smoolzone/ia4ai/main/cli/scripts/install.sh | sh
```

Pin a version, or choose the prefix, through the environment:

```bash
VERSION=v0.1.0 PREFIX="$HOME/.local" sh -c \
  'curl -fsSL https://raw.githubusercontent.com/smoolzone/ia4ai/main/cli/scripts/install.sh | sh'
```

`VERSION` defaults to the latest published release (`v0.1.0` is live); `REPO`
(owner/repo) overrides where the archive is fetched from. **This path needs a
published release** — before `v0.1.0` existed the script stopped with
`could not determine the latest release`. On a private repo with no token it now
says so directly instead of failing on a bare `404`.

**Option 2 — build from source (needs Go 1.22+).**

```bash
make -C cli install                        # -> /usr/local/bin/ia
make -C cli install PREFIX="$HOME/.local"  # no sudo -> ~/.local/bin/ia
make -C cli build                          # just compile -> cli/bin/ia
```

Build through the Makefile, not a bare `go build`: the Makefile is what injects the
`Version`/`Commit`/`BuildDate` metadata, so a bare build reports `dev` forever.
`make -C cli dist` cross-compiles every supported platform into `cli/dist/`.

**Then confirm which binary you got** — `~/.local/bin` is not on `$PATH` by default:

```bash
ia version                          # prints the release tag, or "dev" for a bare build
ia init --base-url http://127.0.0.1:8000
ia login                            # stores a credential, encrypted at rest
ia doctor                           # config, connectivity, role
ia smoke-test                       # api → credential → ingest → read-back → priced
ia case list
ia run show <trace|case>
```

Credentials are AES-256-GCM sealed under a local master key; `--workspace` selects
a stored credential (the server binds tenant to the credential, so it never widens
scope). `ia console` is dispatched against the same server-side allow-list as
the dashboard terminal. The release workflow cross-compiles `linux/darwin × amd64/arm64`
on `v*` tags; the module path lives in one Makefile variable if it needs repointing.

## Deploy with PM2

No containers required — the app is a plain Python package, so PM2 is used only as a
supervisor (autostart, crash-restart, log rotation). A ready-to-edit process file ships
as [`ecosystem.config.js`](ecosystem.config.js).

```bash
npm i -g pm2                        # PM2 is only a process manager, not a runtime
pm2 start ecosystem.config.js       # start the API (ia serve)
pm2 save                            # persist the process list
pm2 startup                         # print/enable the boot hook, then re-run its command
pm2 logs internal-affairs           # tail ./logs/ia.{out,err}.log
pm2 reload internal-affairs         # zero-downtime restart after a deploy
```

Set real values via the environment (or a `.env` next to the process file) rather than
committing secrets: at minimum `IA_EVIDENCE_SIGNING_KEY` **must** be set in production
(`python -c "import secrets; print(secrets.token_hex(32))"`).

**Scaling.** PM2's cluster mode (`-i`) is a Node feature and does not apply to Python.
Keep `instances: 1` with the default SQLite backend (`IA_DB_PATH`) — SQLite is
single-writer, and the tamper-evident evidence/access logs rely on a single store. To
run multiple workers, switch to Postgres (`IA_DB_DSN`, the `[postgres]` extra) and use
uvicorn's own `--workers N` (the commented override in `ecosystem.config.js`). Either
way, terminate TLS at a reverse proxy (nginx/Caddy) in front of the loopback bind.

> `docker-compose.yml` is an **optional dev extra** (a local Postgres for the Postgres
> backend only) — it is not part of the app's runtime and can be ignored entirely.

Endpoints:

- `POST /ingest/trace` — ingest a raw trace (+ optional job) → normalize, redact, log, store.
- `POST /ingest/otlp` — ingest OpenTelemetry Protocol (JSON or Protobuf) trace batches from any OTel source.
  Also served at the standard OTLP/HTTP path `/v1/traces`, so a stock SDK pointed at the base URL works unchanged.
  Store-and-ack: spans are persisted and queued, and the case file is built asynchronously by the worker.
- `GET /case/{case_id}` — fetch a reconstructed case file (read is audit-logged).
- `GET /case/{case_id}/cost-cuts` — recommend cheaper capable models for a job (the core value).
- `GET /diff/{case_a}/{case_b}` — diff two case files (decisions/tools/models/cost).
- `GET /catalog` — the model catalog (providers, tiers, prices, capabilities, notes).
- `POST /catalog/models` — add/update a model (admin) so new models are live immediately.
- `POST /catalog/sync` — sync the catalog from configured feed URLs (admin).
- `GET /audit/evidence-log/verify` — prove the evidence log has not been tampered with.
- `GET /audit/access-log` — who accessed what (immutable, signed access audit).
- `GET /audit/access-log/verify` — prove the access audit log has not been tampered with.
- `GET /trace/{trace_id}` — the span DAG a run *actually executed* (parent/child + timings; read is audit-logged).
- `GET /case/{case_id}/trace` — the same, resolved from a case (tenant-checked via the case).
- `POST /graph/register` — register a static agent blueprint, e.g. `extract_graph_topology(compiled_app)`.
- `GET /graph/{agent_system}` — that blueprint back (the *map*, not the journey; opt-in, 404 is normal).
- `GET /activity` — the pipeline activity feed: the evidence + access logs merged and ordered (admin).
- `GET /activity/stream` — the same feed as Server-Sent Events, for the live dashboard terminal (admin).
- `POST /console` — the dashboard terminal: one line, dispatched against an **allow-list** of read-only
  commands over the same store the API reads (no shell, no subprocess, no writes). Refused with 403 for
  commands your role does not hold; the line is redacted and appended to the access log either way.
- `GET /cost/summary` — aggregate cost: per-model, per-day, per-job rollups + trend drift (viewer).
- `GET /reconcile` — compare metered usage vs the provider bill (Anthropic/OpenAI admin usage APIs) (admin).
- `POST /golden-set/bootstrap` — build a replay golden set from stored production traces (viewer).
- `GET /alerts/history` — recent webhook alert dispatches (viewer).
- `GET /pipeline/stats` — ingestion-worker queue/processing/retention stats (viewer).
- `GET /health` — liveness.

## Test

```bash
uv run pytest
```

The dashboard is inline HTML/JS with no build step, so its **Terminal** and
**Run graph** tabs are covered by a real headless-browser click-through
([`tests/test_ui_clickthrough.py`](tests/test_ui_clickthrough.py)) that also fails
on any JavaScript error the page logs. It skips unless both playwright and a
running server are present:

```bash
uv pip install -e ".[ui]" && playwright install chromium
uv run pytest tests/test_ui_clickthrough.py    # needs `ia serve` running
```

## Project layout

```
internal_affairs/
  schemas/        canonical evidence + case-file models (pydantic)
  security/       redaction + tamper-evident evidence/audit logs
  ingest/         normalization (OTel GenAI semconv) + LangGraph + OTLP adapters
  investigate/    case-file reconstruction + findings + root cause + diffing
  verdict/        cost attribution + "cheaper model?" + policy + evaluation
                  + runner + cascade + judge + cost cutting
  storage/        Store interface: in-memory + SQLite + Postgres
  api/            FastAPI control plane + web UI
examples/         runnable demo + seed data
docs/             architecture + threat model + integration guide
```

Persistence: set `IA_DB_PATH` to a `.db` file for SQLite-backed storage (cases,
traces, evidence log, audit log all persist and stay tamper-evident across restarts).
Leave it empty for in-memory (dev/tests).

Auth/RBAC: disabled by default. Set `IA_AUTH_ENABLED=true` and configure
`IA_AUTH_SECRET` (JWT), `IA_API_KEYS` (`key:role:tenant`), and/or OIDC
(`IA_OIDC_AUDIENCE` + `IA_OIDC_ISSUER`). Endpoints require `investigator`
(ingest), `viewer` (read/diff), or `admin` (audit logs).

## Status

Phase 3 complete (core loop), plus the **verdict-honesty layer**: non-inferiority
three-way verdicts (`viable` / `ruled_out` / `unproven`), governance + latency
gating, verdict lineage, capture→replay (`ia reverify`), and metered verification
cost. An **operations layer** adds an async ingest pipeline (store-and-ack → worker),
provider-bill reconciliation, signed webhook alerting, aggregate cost rollups with
trend drift, and golden-set bootstrap. Deployment is **PM2-based** (`ecosystem.config.js`)
— no containers required. See `docs/architecture.md` for the full roadmap.
