Metadata-Version: 2.4
Name: assurance-budget
Version: 0.1.0
Summary: Where an agent run's budget went, and where the loop went nowhere. Ceilings a caller cannot raise.
License: Apache-2.0
Project-URL: Source, https://github.com/i-ops-hq/assurance-budget
Keywords: ai-agents,observability,cost-control,guardrails,assurance,llm-ops
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: assurance-core>=0.13
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: mypy>=1.0; extra == "dev"
Dynamic: license-file

# assurance-budget

[![tests](https://github.com/i-ops-hq/assurance-budget/actions/workflows/tests.yml/badge.svg)](https://github.com/i-ops-hq/assurance-budget/actions/workflows/tests.yml)
[![PyPI](https://img.shields.io/pypi/v/assurance-budget)](https://pypi.org/project/assurance-budget/)
[![License](https://img.shields.io/badge/license-Apache--2.0-blue)](https://github.com/i-ops-hq/assurance-budget/blob/main/LICENSE)

## Your agent didn't fail. It just kept going.

The expensive runs are rarely the ones that crash. They are the ones that retried the same failing
call fourteen times, or spent nineteen model calls summarising something nobody read, and finished
with a plausible answer and a bill.

Point this at a run log and find out which ones did that.

```bash
pip install assurance-budget
assurance-budget runs.jsonl
```

```
1 of 4 runs hit a limit — 1 was going nowhere first — 1 ran past the clock

  r-002
      Stopped: 3 rounds repeating fetch(url=api/invoices) and failing the same way (timeout)
      with nothing new read and no part of the goal closer.
  r-003
      Stopped after 20 frontier calls
  r-004
      Ran 660s against a 600s cap

  Not tested by this log: iterations, retries. The log carries no events of that kind, so this
  is silence rather than a pass.
```

That last paragraph is the part most tools leave out. **A limit your log cannot exercise has not
passed — it has not been tested**, and reporting those the same way is how a green check comes to
mean nothing.

## A limit a caller can raise is a suggestion

The ceilings are constants in the library, and every budget is clamped to them at construction. Ask
for more and you do not get more:

```bash
assurance-budget runs.jsonl --tool-calls 5000 --json | grep tool_calls
#   "tool_calls": 40
```

It clamps rather than erroring, on purpose. A caller asking for 5000 is expressing a preference the
runtime declines — that is not a reason to abort somebody's task. Lower values pass straight through,
because a caller may always be *more* conservative: that is how a cheap plan or an untrusted worker
gets a shorter leash.

Raising a ceiling is a deliberate edit to a source file, which is the point.

## The log

JSONL, one event per line. Field names are matched loosely — `run`/`run_id`/`session`,
`action`/`tool`/`name`, and so on — so most existing logs work without being rewritten.

```json
{"run": "r-002", "action": "fetch(url=api/invoices)", "error": "timeout", "kind": "tool", "ts": 41.0}
```

`kind` is one of `tool`, `frontier`, `retry`, `iteration` and defaults to `tool`. Anything else is
refused rather than counted as something it isn't.

## As a library — the half that prevents rather than reports

The audit tells you it already happened. This stops it happening:

```python
from assurance_core.run_budget import Budget, ProgressWatch, Progress, Spend

spend = Spend(budget=Budget.allowing(tool_calls=20))
watch = ProgressWatch()

for step in range(100):
    stopped = spend.charge_tool_call()
    if stopped:
        print(stopped.message)
        break
    stalled = watch.observe(Progress(action="fetch(url=api/invoices)", error="timeout"))
    if stalled:
        print(stalled.message)
        break

assert spend.tool_calls <= 20
```

Two different stops, and the difference matters. **`Exhausted` means the budget ran out.
`Stalled` means budget remains and spending it is the mistake** — three rounds repeating the same
action, failing the same way, with nothing new read and no part of the goal closer.

## In a pipeline

```bash
assurance-budget runs.jsonl --fail-on-exhausted
```

| exit | means |
|---|---|
| `0` | audited, and nothing hit a limit or stalled |
| `1` | audited, and `--fail-on-exhausted` found a run that did |
| `2` | **refused** — the log could not be read, so there is no audit |

## Honest limits

- **It reads what your log records.** A run that burned money in a way the log does not mention is
  invisible here, and no amount of analysis fixes that.
- **There is no dollar limit.** Frontier calls are the cost proxy. Prices change per model, per
  provider and per week; a number that goes stale silently is worse than a count that does not
  pretend to be money.
- **Stall detection needs three rounds** and both halves — identical action, error and result, *and*
  flat progress. A repeated action while evidence accumulates is a loop doing work, and stopping
  that would be the bug.

## Where the rules live

`assurance_core.run_budget`, in [`assurance-core`](https://pypi.org/project/assurance-core/) — pure
Python, no dependencies, no model involved in any of it. This package reads logs and calls it.

## Licence

Apache-2.0.
