Metadata-Version: 2.4
Name: tokensbill
Version: 0.7.0
Summary: One line. Every AI token tracked. Captures Anthropic/OpenAI/Gemini token usage from outgoing httpx/requests calls and streams it to TokensBill. Silent-fail by design.
License: MIT
Project-URL: Homepage, https://tokensbill.com
Project-URL: Repository, https://github.com/oxtrys/tokensbill
Keywords: ai,tokens,cost,openai,anthropic,gemini,observability,tokensbill
Requires-Python: >=3.8
Description-Content-Type: text/markdown

# tokensbill

**One line. Every AI token tracked.** Automatically captures Anthropic / OpenAI / Gemini token usage from your app's outgoing AI calls and streams it to your [TokensBill](https://tokensbill.com) dashboard — so you can see exactly where your AI spend goes.

## Install

```bash
pip install tokensbill
```

## Use — one line at startup

```python
import tokensbill; tokensbill.init("tl_live_your_project_key")
```

Put it at the top of your entry file (e.g. `main.py`), before your app starts. That's it — every AI API call your app makes is now tracked automatically.

Get your project key from **TokensBill → your project → Integration**.

## What it does

- Patches `httpx` and `requests` (used by the Anthropic/OpenAI Python SDKs).
- Reads token counts and cost from the AI provider's response **without altering your call**.
- Sends the numbers to TokensBill in a background thread.
- **Silent-fail by design** — it never raises into, blocks, or slows your application. If TokensBill is unreachable, your app is completely unaffected.

## Is the key safe to commit?

The tracking key is a **label, not a password**. It cannot access your AI provider account, read your prompts, or spend money — it only lets your app report usage numbers to your dashboard. You can regenerate it any time from the dashboard.

## Options

```python
import tokensbill
tokensbill.init(
    "tl_live_...",
    environment="production",   # optional - auto-detected from ENVIRONMENT / APP_ENV /
                                # FLASK_ENV / DJANGO_ENV, falling back to "unknown".
                                # Set this only to override.
    ingest_base_url="https://tokensbill.aiappsjunction.com",  # override the endpoint
)
```

## Prove a model swap before you make it

Every model-swap recommendation ends by telling you to check output quality yourself. Shadow replay
does that check — and it tells you what the check will cost before it runs.

```python
import tokensbill

plan = tokensbill.verify("ScoreDocument", "gpt-4.1-mini", samples=50)
print(plan.describe())
plan.start()
```

`verify()` **runs nothing.** It returns a plan:

```
Replay ScoreDocument against gpt-4.1-mini - 50 samples, 100 provider calls
  Estimated cost ~$0.5875 ($0.0118 per sample), based on your measured average of
  2,296 input / 650 output tokens for this function.
  Hard cap $1.00. Nothing runs until you call `.start()`.
```

Only `.start()` arms it. Two steps on purpose: replay spends **your** money, and the first thing you
see about that should not be the invoice.

After it starts, the next `samples` calls to that function have their request bodies held in memory.
Each is re-sent to the candidate model **on your own credentials**, the two answers are compared
**inside your process**, and only a score is reported: `188 of 200 agreed (94%)`. It shows up on the
finding in your dashboard.

**The cost, precisely.** Each sample is *two* calls — the current model and the candidate — because
comparing a fresh candidate answer against a stale cached one would score the passage of time as a
model difference. The estimate is priced from token counts this process has actually measured for
that function. If it hasn't seen the function yet it says **unknown**, never `$0.00`.
`max_cost_usd` is a hard cap: the run stops *before* the sample that would exceed it.

**Want the exact number first?** `dry_run=True` captures real request bodies, prices the run from them,
and sends nothing to the provider and no verdict to TokensBill.

**Your prompts and responses never leave your infrastructure.** The verdict payload is counts, two
model names and a comparison method — the TokensBill API has no field that could receive content.
The answers that disagreed stay on your machine (`plan.result.mismatches`) for you to read.

**Limits, stated plainly.** Comparison is `json` (parsed, property order ignored) or `exact`
(trimmed text). Neither can tell you whether two differently-worded paragraphs mean the same thing,
so this is for structured output. Requests using tools or function calling are **never** replayed —
re-sending one could fire a real action twice — and neither are streaming requests.

## License

MIT
