Metadata-Version: 2.4
Name: llm-workflow-router
Version: 1.1.1
Summary: Deterministic workflow topology enforcement for LLM-powered systems.
Author: Doby Baxter
License-Expression: PolyForm-Noncommercial-1.0.0
Project-URL: Homepage, https://dobybaxter127.gitlab.io/
Project-URL: Repository, https://gitlab.com/dobybaxter127/llm-router
Keywords: llm,agents,workflow,topology,validation,determinism,mcp,langgraph,claude
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Software Development :: Libraries
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: PyYAML>=6.0.1
Provides-Extra: dev
Requires-Dist: pytest>=8.0.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.23; extra == "dev"
Requires-Dist: coverage>=7.0.0; extra == "dev"
Requires-Dist: ruff>=0.6.0; extra == "dev"
Requires-Dist: mypy>=1.8.0; extra == "dev"
Requires-Dist: opentelemetry-api>=1.20; extra == "dev"
Requires-Dist: opentelemetry-sdk>=1.20; extra == "dev"
Requires-Dist: types-PyYAML>=6.0; extra == "dev"
Requires-Dist: jsonschema>=4.18; extra == "dev"
Provides-Extra: otel
Requires-Dist: opentelemetry-api>=1.20; extra == "otel"
Provides-Extra: security
Requires-Dist: opentelemetry-api>=1.20; extra == "security"
Provides-Extra: openai-agents
Requires-Dist: openai-agents>=0.18; extra == "openai-agents"
Provides-Extra: claude-agent-sdk
Requires-Dist: claude-agent-sdk>=0.1; extra == "claude-agent-sdk"
Provides-Extra: langgraph
Requires-Dist: langgraph>=1.0; extra == "langgraph"
Provides-Extra: mcp
Requires-Dist: mcp>=1.0; extra == "mcp"
Dynamic: license-file

# LLM Workflow Router

Deterministic workflow topology enforcement for LLM-powered systems.

LLM Workflow Router is a stateless middleware engine designed to enforce explicit execution topology in AI systems that rely on large language models. It evaluates structured interaction metadata against strictly declared workflow rules and returns a terminal state.

It controls structure — not content.

---

## Overview

Modern LLM-driven systems frequently suffer from:

- Recursive tool invocation loops  
- Circular container transitions  
- Cross-container contamination  
- Implicit fallback behavior  
- Unbounded workflow escalation  
- Inconsistent refusal logic  

Most mitigation strategies limit volume (timeouts, max tool calls, retries).  
LLM Workflow Router enforces topology explicitly.

The engine evaluates structured interaction metadata and returns one of three terminal states:

- PROCEED  
- REFUSE  
- PAUSE  

No content inspection.  
No moderation.  
No orchestration.  
No mutation of input.  

Only structural enforcement.

---

## Core Design Principles

- Deterministic evaluation  
- Stateless per evaluation  
- Metadata-only inspection  
- Explicit transitions only  
- Static configuration (v1)  
- Strict validation at load time  
- No silent fallback  
- Host application retains execution control  

Same input → same output.

---

## Architecture

Application  
↓  
WorkflowEngine.evaluate(metadata)  
↓  
[ PROCEED | REFUSE | PAUSE ]  
↓  
Application decides next action  

The router does not:

- Execute tools  
- Retry calls  
- Modify prompts  
- Orchestrate sessions  

It enforces topology and returns a decision.

---

## Configuration

Workflow rules are defined using static YAML configuration.

```yaml
entrypoints:
  - entry

containers:
  entry:
    allow_transitions:
      - support
      - REFUSE
    allow_reentry: false
    max_invocations: 2

  support:
    allow_transitions:
      - faq
      - REFUSE
    allow_reentry: false
    max_invocations: 3

  faq:
    allow_transitions:
      - REFUSE
    allow_reentry: true
    max_invocations: 5
```

Each container defines:

- Explicit allowed transitions  
- Whether re-entry is permitted  
- Maximum invocation depth  

Implicit transitions are not allowed.

### Entrypoints

The optional top-level `entrypoints` list declares the authoritative root
containers. A workflow may have several roots — a support flow, a sales flow,
and an FAQ flow can each start in their own container — so `entrypoints` is a
list. When declared, any indegree-0 container that is *not* listed is treated
as an orphan and rejected. If `entrypoints` is omitted, roots are inferred from
graph shape (indegree 0) for backward compatibility, and an ambiguity warning
is raised when more than one root is inferred. See
`examples/multi_entry_config.yaml`.

### Transition budget

`max_total_transitions` caps a whole run, not just one container:

```yaml
max_total_transitions: 6
```

Once `transition_history` holds that many containers, any further
container-to-container transition is refused with
`MAX_TOTAL_TRANSITIONS_EXCEEDED`. Terminal targets (`PROCEED`, `REFUSE`,
`PAUSE`) are always allowed, so a run that spent its budget can still end
cleanly. Without it, a cycle of re-entrant containers is bounded only by each
container's own `max_invocations`.

### Approval gates

Mark a container `requires_approval: true` and entering it needs a human (or
any approver you choose):

```yaml
  refunds:
    requires_approval: true
    allow_transitions: [PROCEED, REFUSE]
```

A structurally valid transition into `refunds` returns `PAUSE` with reason
`APPROVAL_REQUIRED` until the host re-evaluates the same metadata with
`approval_granted: true`. The engine stays stateless: approval is an input, not
something it remembers. See `examples/approval_config.yaml`.

### Strict parsing and editor support

Unknown keys are rejected at load time, so a typo such as `max_invocation: 5`
fails loudly instead of silently falling back to the default. For
autocompletion and inline errors while editing, export the JSON Schema and
point your editor at it:

```bash
llm-router schema > workflow.schema.json
```

```yaml
# yaml-language-server: $schema=./workflow.schema.json
```

---

## Topology Validation

At configuration load time, the router performs strict validation:

- Invalid transition targets  
- Unknown containers  
- Unknown declared entrypoints  
- Dead-end containers  
- Missing entry points  
- Orphan containers (indegree 0, not a declared entrypoint)  
- Unreachable containers  
- Self-transition contradictions  
- Cycle detection  
- Re-entry safety enforcement  
- Containers beyond the transition budget  
- Approval gates on entrypoints (where they have no effect)  

Configuration errors raise exceptions immediately.

Fail loudly at load time.  
Never fail silently at runtime.

---

## Runtime Evaluation

The engine evaluates an immutable metadata structure:

```python
InteractionMetadata:
    container: str
    previous_state: InteractionState
    transition_history: List[str]
    invocation_depth: Dict[str, int]
    requested_action: str
    trace_id: Optional[str]
    approval_granted: bool = False
```

Returns:

```python
EvaluationResult:
    state: InteractionState
    container: str
    reason: Optional[ReasonCode]
    trace_id: Optional[str]
    allowed_transitions: Tuple[str, ...]
```

No exceptions during normal evaluation.  
Only terminal states are returned.

Every `REFUSE` says what *would* have been accepted:
`allowed_transitions` lists the targets the same metadata could have requested
(empty when the container itself is exhausted). An agent that took a wrong turn
can recover without guessing, and `engine.allowed_targets(metadata)` gives the
same list before anything is attempted.

Load a config from Python with the same strict checks the CLI uses:

```python
from router import WorkflowEngine, load_config

engine = WorkflowEngine(load_config("config.yaml"))
```

---

## CLI Usage

Install:

```bash
pip install llm-workflow-router
```

Validate configuration:

```bash
llm-router validate config.yaml
```

Analyze topology:

```bash
llm-router analyze config.yaml
```

Evaluate metadata:

```bash
llm-router run --config config.yaml --metadata metadata.json
```

---

Print the JSON Schema for configs:

```bash
llm-router schema
```

Review a config change. Every change is tagged `WIDENS`, `NARROWS` or
`NEUTRAL`, and newly reachable containers are listed. Classification is
conservative, so `--fail-on-widen` works as a CI gate (exit code 1):

```bash
llm-router diff old.yaml new.yaml --fail-on-widen
```

Replay recorded traffic against a new config and see every decision that would
now come out differently. A trace is the JSON Lines log that
`run --verbose` writes to stderr (or `EvaluationLogEvent.from_result(...)` from
your own code). Add `--secure` to replay through the Security Layer too:

```bash
llm-router run --config config.yaml --metadata md.json --verbose 2>> trace.jsonl
llm-router replay trace.jsonl --config new.yaml --fail-on-change
```

Exit codes: `0` success, `1` a `diff` or `replay` gate tripped, `2` the input
could not be read or validated.

Render the topology as a diagram:

```bash
llm-router graph config.yaml                # Mermaid (paste into any Mermaid renderer)
llm-router graph config.yaml --format dot   # Graphviz DOT
```

Like `analyze`, `graph` renders broken topologies too — unknown transition
targets are drawn and flagged, which is exactly what you want while debugging
a config.

---

## Session Layer (optional)

The engine is stateless by design: every `evaluate()` receives a complete
metadata snapshot. If you'd rather not do that bookkeeping yourself,
`WorkflowSession` does it for you — and only advances on `PROCEED`:

```python
from router import WorkflowEngine, WorkflowSession

engine = WorkflowEngine(cfg)
session = WorkflowSession(engine, entry="entry", trace_id="req-42")

result = session.request("support")   # entry -> support
result = session.request("faq")       # support -> faq
result = session.request("REFUSE")    # terminal; session closes

session.history           # ("entry", "support")
session.invocation_depth  # {"entry": 1, "support": 1, "faq": 1}
```

A `REFUSE` closes the session. A `PAUSE` suspends it until `resume()`.
`session.snapshot(target)` exposes the exact metadata the next request would
evaluate, so the session is fully auditable and you can drop down to raw
`engine.evaluate()` at any time. The engine itself remains pure and stateless.

At an approval gate the session pauses and holds the pending target:

```python
result = session.request("refunds")   # PAUSE, reason APPROVAL_REQUIRED
session.pending_approval               # "refunds"
result = session.approve()             # re-evaluates with approval; PROCEED
# or session.resume() to decline and stay where you are
```

The gate itself does not consume an invocation; the approved retry does.

---

## OpenAI Agents SDK Integration

Agents are containers. Handoffs are transitions. Attach a `TopologyGuard` to
a run and every handoff is structurally evaluated **before** the next agent
executes — a refused handoff raises `TopologyViolation` and aborts the run
loudly instead of letting the agent graph wander:

```bash
pip install "llm-workflow-router[openai-agents]"
```

```python
from agents import Agent, Runner
from router import WorkflowEngine
from router.integrations.openai_agents import TopologyGuard, TopologyViolation

guard = TopologyGuard(engine, entry="triage", trace_id="run-001")

try:
    result = await Runner.run(triage_agent, "I was double-charged.", hooks=guard)
except TopologyViolation as violation:
    print("Refused:", violation.result.reason)

# Full structural audit trail of the run:
print(guard.session.history, guard.session.invocation_depth)
```

By default an agent's `name` is its container name; pass `container_for=` to
map differently. For handoffs into `requires_approval` containers, pass
`approver=`, a sync or async `(source, target) -> bool`. Without one, or when it
returns `False`, the handoff raises `TopologyViolation`. The guard is content-blind — it never reads prompts,
messages, or tool arguments. See
[`examples/openai_agents_example.py`](examples/openai_agents_example.py) for a
complete runnable triage → billing → refunds system.

---

## Tool Gating: Claude Agent SDK and MCP

For single-agent systems the risky structure is usually the *sequence of tool
calls*, not handoffs. Here **tools are containers and each tool call is a
transition** from the tool called before it:

```yaml
entrypoints: [start]
containers:
  start:      { allow_transitions: [search], max_invocations: 3 }
  search:     { allow_transitions: [search, summarize], allow_reentry: true, max_invocations: 5 }
  summarize:  { allow_transitions: [send_email], max_invocations: 2 }
  send_email: { allow_transitions: [PROCEED], requires_approval: true }
```

"Never send an email before summarizing, and never without a human" is now a
contract rather than a prompt instruction. A refused call does **not** crash
the run: the model is told which tools it may call next and can recover, and
repeated bad calls exhaust `max_invocations`, so recovery is bounded. Every
tool must be declared or listed in `passthrough`, and a tool whose name maps to
`PROCEED`, `REFUSE` or `PAUSE` is refused, since those are workflow states;
nothing is waved through silently. Calls in one conversation are decided one at
a time, so a call waiting on an approver holds the calls behind it until the
answer is in. The gate reads tool names only, never arguments or results.

**Claude Agent SDK** (`pip install "llm-workflow-router[claude-agent-sdk]"`):

```python
from claude_agent_sdk import ClaudeAgentOptions, query
from router.integrations.claude_agent_sdk import TopologyHooks

topology = TopologyHooks(engine, entry="start", passthrough={"Read", "Glob"})
options = ClaudeAgentOptions(hooks=topology.hooks())
```

Refused calls are denied through `PreToolUse` with a reason naming the
allowed tools. Permitted calls return no decision, so your normal permission
rules still apply. Each sub-agent gets its own topology session (keyed by
`agent_id`, with `entry_for=` to start sub-agent types elsewhere), so parallel
sub-agents never tangle. At an approval gate, pass `approver=` to decide
in-process, or leave it out and the gate defers to the SDK's own permission
prompt.

**MCP** (`pip install "llm-workflow-router[mcp]"`):

```python
from router.integrations.mcp import GatedClientSession

session = GatedClientSession(client_session, engine, entry="start")
result = await session.call_tool("search", {"q": "..."})
```

A refused call never reaches the server. It comes back as a normal
`CallToolResult` with `isError` set, which is how MCP reports tool failures to
a model, so your agent loop needs no special handling. Everything besides
`call_tool` is passed through to the wrapped session.

Both adapters are thin layers over `router.integrations.tool_gate.ToolGate`,
which you can call directly from any other framework. The gate keeps a session
per conversation key until told otherwise, so a long-running host should call
`gate.forget(key)` when a conversation ends (each adapter exposes its gate as
`.gate`).

---

## LangGraph Integration

**Nodes are containers; running a node is a transition.** LangGraph's edges
say where a graph *may* go. The guard enforces your declared topology
independently, so a routing function or `Command(goto=...)` that strays
outside the contract raises `TopologyViolation` instead of running the node:

```bash
pip install "llm-workflow-router[langgraph]"
```

```python
from router.integrations.langgraph import TopologyGuard

guard = TopologyGuard(engine, entry="triage")
builder.add_node("triage", guard.node("triage", triage))
builder.add_node("refunds", guard.node("refunds", refunds))
```

Approval gates use LangGraph's own human-in-the-loop: the guard calls
`interrupt()` with `{"type": "wfrouter.approval_required", "source": ...,
"target": ...}`, and you resume with `Command(resume=True)` to approve or
`Command(resume=False)` to decline (a checkpointer is required, as for any
interrupt). A gated node can still call `interrupt()` itself, and its own
questions receive their own answers.

Sessions are kept per `thread_id`, in the guard's memory, and one session
covers one run of the graph. LangGraph starts every new invocation with fresh
input at `START`, so on a thread that has run before, call `start_run` first.
Resuming with `Command(resume=...)` continues the current run and needs no call:

```python
config = {"configurable": {"thread_id": "chat-1"}}
guard.start_run("chat-1")  # new input on a thread that has run before
graph.invoke(new_input, config)
```

Without it, the entry node is refused as a transition from wherever the last
run ended, and the error names the call to make. Call `guard.forget(thread_id)`
when a conversation ends, so a long-running host does not keep every thread in
memory. The guard follows one node at a time, so graphs that fan out to
parallel branches in a single step are outside its model.

---

## Why not just LangGraph (or my orchestrator's built-in graph)?

Orchestrators *describe* structure. This engine *enforces* it — as a separate,
framework-agnostic layer with properties orchestrators don't give you:

- **Independent enforcement.** The topology lives outside your agent
  framework, so a prompt-induced detour, a buggy handoff, or a framework
  upgrade can't silently widen what's reachable. The declared graph is a
  contract, and violations fail loudly with structured reason codes.
- **Framework-agnostic.** The same YAML config governs an OpenAI Agents SDK
  app today and whatever you migrate to next year. Adapters are thin; the
  contract is portable.
- **Auditable determinism.** Same metadata + same config → same decision,
  every time. Combined with the `wfrouter.*` OpenTelemetry conventions, you
  get compliance-grade answers to "why was this transition refused?" — a
  reason code, not a vibe.
- **Load-time topology analysis.** Cycles, orphans, dead ends, and
  unreachable states are caught before anything runs — the kind of static
  validation industrial control systems have had for decades and agent
  frameworks mostly don't.

If you're happy inside one orchestrator and don't need independent structural
guarantees, its built-in graph may be enough. This tool exists for when
"probably follows the graph" isn't good enough.

---

## Intended Audience

- AI SaaS developers  
- Internal LLM tooling teams  
- Agent orchestration builders  
- Platform engineering teams  
- Infrastructure-focused AI developers  

Not intended for content moderation or prompt filtering.

---

## Observability

Optional OpenTelemetry instrumentation is provided under a dedicated,
versioned namespace (`wfrouter.*`) that this project owns — it is deliberately
independent of the upstream `gen_ai.*` conventions, which assume a model at the
center of every span and do not fit a content-blind topology engine.

Install the extra:

```bash
pip install "llm-workflow-router[otel]"
```

Wrap evaluation:

```python
from router.observability.otel import traced_evaluate
result = traced_evaluate(engine, metadata)
```

If `opentelemetry-api` is not installed, instrumentation degrades to a no-op
and the engine behaves identically. A `PROCEED`, `REFUSE`, or `PAUSE` outcome
is a successful decision (span status OK); only genuine failures are errors.
See `OBSERVABILITY.md` for the full attribute and span conventions.

---

## Security Layer (`router.security`) — commercial

An optional layer that secures the *structure* of a workflow: it enforces which
trust levels may reach which side-effecting capabilities, emits
`wfrouter.security.*` telemetry, and keeps a tamper-evident audit trail. Same
content-blind, deterministic principles as the core.

Declare trust posture per container, then check it:

```yaml
containers:
  intake:      { trust: UNTRUSTED, allow_transitions: [guard] }
  guard:       { sanitizer: true,  allow_transitions: [tool] }
  tool:        { capability: TOOL_EXEC, allow_transitions: [PROCEED] }
```

```bash
llm-router secure config.yaml        # reports trust-boundary issues
```

```python
from router.engine import WorkflowEngine
from router.security import SecurityEngine, AuditLog

secure = SecurityEngine(WorkflowEngine(cfg), audit_log=AuditLog())
result = secure.evaluate(metadata)   # REFUSE if a boundary is crossed
assert secure.audit_log.verify()
```

The core property: *untrusted input must pass a sanitizer before it can reach a
side-effecting capability.* See [SECURITY.md](SECURITY.md) for the full model,
findings, and telemetry conventions.

> **Licensing:** the whole project is **PolyForm Noncommercial 1.0.0** — free to
> inspect, run, and use for any noncommercial purpose, **commercial use requires a
> paid license**. See [LICENSING.md](LICENSING.md) and [PRICING.md](PRICING.md).

---

# License

**Source-available under [PolyForm Noncommercial 1.0.0](LICENSE).** The whole
project — core engine and Security Layer alike — is free to inspect, run, modify,
and use for any **noncommercial** purpose: personal projects, research, education,
non-profits, and evaluation.

**Commercial use requires a paid license.** That means using any part of it in a
product or service you sell or host, inside a for-profit company's production or
internal systems, or in paid client work. See [LICENSING.md](LICENSING.md),
[COMMERCIAL-LICENSE.md](COMMERCIAL-LICENSE.md), and [PRICING.md](PRICING.md).

If you deploy this in production, a note about your use case is always
appreciated (and helps prioritize the roadmap) — but never required.

---

## Version

Current version: 1.1.1

- Static configuration model
- Explicit multi-entrypoint declaration (with inferred fallback)
- Whole-run transition budget and human approval gates
- Refusals that name the allowed next steps
- Optional stateful `WorkflowSession` convenience layer
- Adapters: OpenAI Agents SDK, Claude Agent SDK, LangGraph, MCP
- Mermaid / DOT topology export (`llm-router graph`)
- Change review: `llm-router diff` and `llm-router replay`
- JSON Schema for configs (`llm-router schema`)
- **Security Layer: trust-boundary enforcement, `wfrouter.security.*`
  telemetry, tamper-evident audit trail (commercial)**
- Licensed: PolyForm Noncommercial 1.0.0 (noncommercial free · commercial licensed)

No inheritance. No dynamic rule composition.

Future versions may extend topology modeling capabilities.

---

## Author

Doby Baxter  
Software systems developer focused on deterministic infrastructure and human-centered tooling.
