Metadata-Version: 2.5
Name: agent-budget-guard-py
Version: 0.4.0
Summary: Budget caps, runaway-loop detection and circuit breakers for AI agents. Zero dependencies, framework agnostic.
Project-URL: Homepage, https://github.com/yaoyuxiang-gnn/agent-guard
Project-URL: Repository, https://github.com/yaoyuxiang-gnn/agent-guard
Project-URL: Issues, https://github.com/yaoyuxiang-gnn/agent-guard/issues
Project-URL: Changelog, https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/CHANGELOG.md
Project-URL: Roadmap, https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/ROADMAP.md
Author-email: Yuxiang Yao <yaoyuxiangyyx@gmail.com>
License: MIT License
        
        Copyright (c) 2026 Yuxiang Yao
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
License-File: LICENSE
Keywords: agent,agents,ai,anthropic,budget,circuit-breaker,cost,guardrails,llm,loop-detection,observability,openai,token-usage
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Typing :: Typed
Requires-Python: >=3.10
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.34; extra == 'anthropic'
Provides-Extra: dev
Requires-Dist: mypy>=1.11; extra == 'dev'
Requires-Dist: pytest-cov>=5.0; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.16; extra == 'dev'
Provides-Extra: langgraph
Requires-Dist: langgraph>=0.2; extra == 'langgraph'
Provides-Extra: openai
Requires-Dist: openai>=1.40; extra == 'openai'
Description-Content-Type: text/markdown

<h1 align="center">agent-guard</h1>

<p align="center"><b>Stop your agent before it burns $400 overnight</b><br>
budget caps, loop detection and circuit breakers for AI agents<br>
zero dependencies · no proxy · nothing leaves your process</p>

<p align="center"><img src="https://raw.githubusercontent.com/yaoyuxiang-gnn/agent-guard/main/docs/demo.svg" width="600" alt="A terminal running examples/basic.py: an agent with a $0.05 budget is stopped on its third call, and the report shows the budget bar at 114%"></p>

An agent that stops making progress does not stop running. It calls the same tool
with the same arguments, or bounces between two tools forever, buying a fresh
context window on every pass. `agent-guard` is the part of your loop that notices —
and the receipt that tells you what it cost.

```bash
pip install agent-budget-guard-py
```

```python
from agentguard import Guard

guard = Guard(max_usd=1.00, max_steps=25)

with guard:
    while True:
        with guard.step() as step:
            response = client.chat.completions.create(...)
            step.record(response)                     # tokens and cost, extracted
            with step.tool("search", {"q": query}):   # fingerprinted for loop detection
                results = search(query)

print(guard.report())
```

Three things happen on their own: `step.record(response)` pulls tokens and the model
out of *any* SDK response, `step.tool(...)` fingerprints the call so a repeat is
caught **before** it runs a fourth time, and `guard.report()` prints the receipt.
Python 3.10+, no runtime dependencies — not even a provider SDK.

> **Jump to:** [See it work](#see-it-work) · [What it catches](#what-it-catches) ·
> [Wired into your stack](#wired-into-your-stack) · [Models it doesn't know](#models-it-doesnt-know) ·
> [Running it in CI](#running-it-in-ci) · [Honest answers](#honest-answers) ·
> [API reference](https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/docs/API.md) ·
> [All the details](https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/docs/DETAILS.md)

---

## The problem

A budget check that runs *after* the call can only tell you what you already spent.
By the time a cap notices a stuck agent, the money is gone — and the usual failure is
not a crash, it is a loop that looks busy.

Two things stop it, and they stop different things:

| | Catches | When it fires |
|---|---|---|
| **A ceiling** — `max_usd`, `max_tokens`, `max_steps`, `max_seconds` | spend that keeps climbing | after the call that crossed it |
| **A loop detector** — four of them, on by default | an agent repeating itself | **before** the repeat runs |

`agent-guard` does both, in your process, using the standard library. No proxy, no
server, no account, and nothing is sent anywhere.

## See it work

A research agent asked the same question three times. The third identical
`search_web` call is where it stopped — not because the budget ran out, but because
the call was identical to the two before it. This is the whole library in one run:

```python
guard = Guard(max_usd=5.00, name="research-agent")
with guard.step(tag="search") as step:
    step.record("gpt-4o", input_tokens=8_000, output_tokens=400)
    with step.tool("search_web", {"query": "weather in oslo"}):
        results = search("weather in oslo")        # 3rd time: stopped here
```

```text
turn 1: searched, spent $0.0240
turn 2: searched, spent $0.0480

LoopDetected: Loop detected [repeat]: the same call appeared 3 times in the
last 3 steps: search_web({"query":"weather in oslo"}).
```

```text
agentguard  research-agent
================================================================
  wall time   0ms             steps     3
  llm calls   3               tokens    25,200  (in 24,000 / out 1,200)

  limits
    budget   $0.072 / $5                   1.4%  [................]

  by model
    gpt-4o   3 calls      $0.072      24,000 in / 1,200 out

  tripped: loop [repeat] the same call appeared 3 times in the last 3 steps: ...
```

The difference between that and an overnight bill is $0.07. Both output blocks are
real — the exception line is wrapped to fit here and the report is truncated the same
way. Two examples run offline and show the same machinery:

```bash
python examples/loop_detection.py   # all four detectors, plus a healthy run that must not trip
python examples/basic.py            # a budget cap instead: stopped at $0.057 of a $0.05 limit
```

## What it catches

Every limit is optional and independent. A guard with no limits still detects loops
and still produces a report.

| Limit | Trips when |
|---|---|
| `max_usd=1.00` | Known spend exceeds one dollar |
| `max_tokens=500_000` | Input + output tokens exceed the allowance |
| `max_steps=25` | A 26th `guard.step()` is opened |
| `max_seconds=300` | Wall-clock time since the guard was created |
| `scoped_budgets={"tool:search": 1.0}` | One tool, or one tag, exceeds its own share |

A run-level ceiling can only tell you the run got expensive. It cannot tell you
*which part* of it did, so one tool in a retry storm can spend the whole allowance
before `max_usd` notices — and the answer to "what ate the budget?" arrives in a
report you read afterwards. A scoped budget answers it while it is happening, by
naming the tool:

```text
  limits
    budget     $0.0125 / $1                  1.2%  [................]
  ! tool:fetch $0.0125 / $0.012            104.2%  [################]
```

The four detectors, and what each one is actually for:

| Detector | Catches | Fires on |
|---|---|---|
| `RepeatDetector` | The same call, identical arguments | 3rd identical call |
| `CycleDetector` | `A, B, A, B` — two tools bouncing forever | 2nd full pass |
| `SimilarityDetector` | Paraphrasing: `search("python asyncio")` → `search("python asyncio ")` | 4th near-identical call |
| `NoProgressDetector` | A progress marker that never moves | 6th unchanged marker |

Every verdict is a sentence, because "why did you kill my agent?" is the first
question anyone asks. These are real, from `examples/loop_detection.py`:

```text
exact repeat                     -> repeat
                                    the same call appeared 3 times in the last 3 steps: search_web({"query":"weather in oslo"})
two-step ping-pong               -> cycle
                                    a 2-step pattern repeated 2 times: read_file({"path":"app.py"}) -> write_file(...)
paraphrased calls                -> similarity
                                    4 near-identical calls (>= 95% similar) in the last 4 steps: search_web(...)
no progress                      -> no-progress
                                    the progress marker did not change for 6 consecutive observations: {"rows_written":0}

healthy varied work              -> clean, as it should be
```

That last line matters as much as the others: a detector that fires on healthy work
is a detector you turn off. Each threshold is tuned so its scenario is reachable and
nothing else is, and every one is adjustable per guard.

## Wired into your stack

Three depths. Nothing is required beyond the first.

**1. Manual** — works with anything, including a `while` loop you wrote by hand:

```python
guard = Guard(max_usd=1.0)
guard.record("gpt-4o", input_tokens=1200, output_tokens=300)
guard.check()
```

**2. Structured** — one step per iteration, tools fingerprinted for you. This is the
Quickstart above.

**3. Wrapped client** — every call accounted with no call-site changes:

```python
from openai import OpenAI
from agentguard.adapters.openai import guard_openai

client = guard_openai(OpenAI(), max_usd=1.0, max_steps=25)
response = client.chat.completions.create(...)   # recorded automatically
```

Adapters are pure duck-typing — they never import a provider SDK — so the same
wrapper covers **Anthropic, LiteLLM, OpenRouter, vLLM, Together, Groq and Azure
OpenAI**:

```python
from agentguard.adapters.anthropic import guard_anthropic

client = guard_anthropic(Anthropic(), max_usd=2.0)
```

**Streaming works.** With `stream=True` the response is wrapped, chunks pass through
untouched, and usage is recorded once when the stream drains. Anthropic's
`messages.stream()` and async clients (`async for`) are covered the same way.

**LangGraph** takes one callback handler:

```python
from agentguard.integrations.langgraph import guard_langgraph

handler = guard_langgraph(Guard(max_usd=1.0, max_steps=25))
graph.invoke(inputs, config={"callbacks": [handler]})
```

**Or a decorator**, for a request handler rather than a loop:

```python
from agentguard import current_guard, guarded

@guarded(max_usd=0.50, max_steps=20)
def summarise(url: str) -> str:
    guard = current_guard()
    ...
```

`@guarded(max_usd=...)` creates a **fresh guard per call** — the right default when
one caller exhausting a budget must not stop the next. Pass `guard=` a shared guard
when spend should accumulate across calls.

The decorator covers `async def` too: the guard stays entered across every `await`
rather than only until the coroutine object is created, so `current_guard()` is live
inside an async body — and inside an async generator for the whole iteration.

**Fanning work out across threads?** No thread inherits a `contextvars` context, so a
call made on a pool worker is still *counted* — the money is never lost — but belongs
to no step and no tool, and a `scoped_budgets` cap on that tool never fires. Two ways
to carry the attribution over, both captured where you call them:

```python
with guard.step(tag="fan-out"), guard.tool("fetch"):
    results = list(pool.map(guard.bind(fetch), urls))   # per call
    # or: with guard.context():  ...                    # per block, inside the worker
```

**Refuse a call before paying for it.** A post-hoc check can only report overspend;
`preflight()` refuses a call whose worst case will not fit in what is left:

```python
client = guard_openai(OpenAI(), max_usd=0.05, preflight=True)
# BudgetExceeded: refused before spending: a gpt-4o call could reach $0.6,
# over the $0.05 limit (already spent $0)
```

**Resume without refilling the budget.** An agent that checkpoints its own state can
checkpoint its spend too, so a restart does not hand it a fresh cap:

```python
write_checkpoint({"cursor": 41, "guard": guard.snapshot()})

# later, in a new process
guard = Guard.from_snapshot(read_checkpoint()["guard"], max_usd=5.0)
guard.remaining_usd      # what is actually left, not the full budget
```

Three decisions worth knowing, all in
[the details](https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/docs/DETAILS.md#checkpointing):
detector windows survive a checkpoint (a loop that spans one is still a loop);
wall-clock time does not (`max_seconds` caps *this process*, and restoring elapsed
time would trip on time you never spent); and a snapshot it cannot read exactly is
**refused** rather than half-applied.

## Models it doesn't know

The bundled table carries **119 models** — current families from OpenAI, Anthropic,
Google, xAI, DeepSeek, Mistral and Qwen, dated `2026-09` in `PRICING_AS_OF`. It
cannot know your fine-tune, your gateway's aliases, or your negotiated rate:

```bash
agentguard config set my-finetune-v3 3 12 --cached 0.3
agentguard config alias acme/fast claude-3-5-haiku
agentguard config disable gpt-4          # do not trust this bundled price
```

**A model it has never heard of is never billed at a guessed rate.** It is counted
as *unpriced*, kept out of the budget arithmetic, and reported loudly — because a
safety tool that silently assumes `$0.00` is worse than no safety tool at all:

```text
  by model
    gpt-4o            1 call       $2.52   1,000,000 in / 2,000 out
    acme-rerank-v3    1 call    unpriced      40,000 in / 0 out

  ! 1 call(s) had no known price and are excluded from the budget:
      acme-rerank-v3
    Price them with `agentguard config set <model> <input> <output>`,
    or pass Guard(pricing={...}) in code.
```

A model *disabled* by `disable` becomes unpriced the same way, so the excluded spend
is visible instead of quietly billed at a number you rejected. Everything is
overridable in code, and code always beats a file.

## Running it in CI

`guard.save()` in the worker, `agentguard report` in CI — the reading process needs
nothing installed:

```bash
agentguard report run.json           # render a report saved by guard.save(...)
agentguard report run.json --json
agentguard pricing                   # effective table, with a source per model
agentguard pricing gpt-4o
agentguard pricing --update          # refresh prices from a public catalogue
agentguard pricing --status          # is a snapshot in effect, and from where
agentguard config path               # where config is read from, what is ignored
```

## Honest answers

**Zero dependencies, and it stays that way.** No `pydantic`, no `httpx`, no provider
SDK — standard library only. CI fails the build if a runtime dependency ever appears
in the wheel, so it can be dropped into a stack that vendors its dependencies, pinned
to an old Python, or shipped inside a Lambda without touching a lockfile.

**A price table goes stale.** This one is dated, and a provider can change a rate or
retire a model the day after. Retired models keep their last published price rather
than being dropped, because removing a name silently turns every call to it unpriced.
Anything you bill on should be verified, and anything you depend on should be
configured. When you would rather not wait for a release, refresh it on purpose:

```bash
agentguard pricing --update     # one opt-in download, checksummed and cached
```

That is the only command in this library that touches the network. Nothing fetches
prices at import, on a timer, or in the background — and the result merges
*underneath* your config, so a public catalogue can never override a rate you set
deliberately. A snapshot that has been edited or truncated fails its checksum and is
refused rather than billed. See [the details](https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/docs/DETAILS.md#refreshing-the-table).

**What it is not.** Not an observability platform (nothing is sent anywhere, there is
no server and no background thread), not a proxy (it cannot see traffic it was not
told about), not a tokenizer (pre-flight estimates are heuristic), and not a
substitute for provider-side spend limits — use both. agent-guard stops *your* loop;
the provider's limit is what saves you when your process dies with a request already
in flight.

**Limitations, stated plainly.** A response that reports no usage cannot be priced: it
warns once and counts as unpriced rather than inventing a number. A stream is priced
when drained, so for OpenAI-compatible streams pass
`stream_options={"include_usage": True}`. Pre-flight input counts are estimated from
the serialised prompt.

## Documentation

| | |
|---|---|
| [docs/API.md](https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/docs/API.md) | Every public name: `Guard` and `Step`, recording, loop detection, accounting, pricing, config, exceptions, adapters, decorators, the CLI |
| [docs/DETAILS.md](https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/docs/DETAILS.md) | The reasoning: every detector and its tuning, the pricing config's trust model, the checkpoint format, design principles |
| [examples/](https://github.com/yaoyuxiang-gnn/agent-guard/tree/main/examples) | Ten runnable programs — budget cap, all four detectors, wrapped client, streaming, LangGraph, custom models, checkpointing, scoped budgets, a price refresh, and a realistic report |
| [CHANGELOG.md](https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/CHANGELOG.md) | Release history |
| [ROADMAP.md](https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/ROADMAP.md) | What is planned next |
| [CONTRIBUTING.md](https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/CONTRIBUTING.md) | The four constraints the library is built to |

Every docstring example in the package runs as a test, so the documentation cannot
drift from the behaviour — and the API reference is checked against the library the
same way: its signature blocks, its export table and its exception tree are all
asserted in `tests/test_api_reference.py`. `python -m unittest discover -s tests -t .`
runs the whole suite — 746 tests, no network, no fixtures.

## License

MIT — see [LICENSE](https://github.com/yaoyuxiang-gnn/agent-guard/blob/main/LICENSE).
