Metadata-Version: 2.4
Name: vantage-litellm-callback
Version: 0.0.4
Summary: Durable Vantage lifecycle callback for LiteLLM
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: litellm
Requires-Dist: pydantic>=2

# Vantage LiteLLM Callback

## Required Vantage collector

The callback works with the Vantage LiteLLM collector. To set up the Vantage
LiteLLM integration, install both components:

1. Run the collector beside each LiteLLM proxy. The collector stores usage
   events and uploads compacted batches to Vantage.
2. Install this callback in the LiteLLM Python environment. The callback sends
   usage events to the local collector over its Unix socket.

The callback cannot upload usage to Vantage by itself. The collector cannot
collect LiteLLM usage without the callback. See the collector repository README
for deployment and configuration instructions.

## Install

Install a released version with:

```sh
pip install vantage-litellm-callback
```

To work from a source checkout, install this directory with `pip install .`.
Configure LiteLLM with:

```yaml
litellm_settings:
  callbacks:
    - vantage_callback.callback_instance
```

See the repository README for collector setup and lifecycle billing semantics.

The callback is fail-open: collector outages, acknowledgement failures, malformed
responses, queue pressure, and projection errors never block or fail a provider
request. Usage events are placed into a bounded in-process queue and delivered by
a background worker. When delivery is unhealthy, the worker opens a circuit,
aggregates bounded loss information, and periodically probes the collector with a
versioned `delivery_gap` control frame. Normal delivery resumes only after that
frame receives a durable acknowledgement.

Configuration:

- `VANTAGE_COLLECTOR_DELIVERY_QUEUE_SIZE` (default `1024`)
- `VANTAGE_COLLECTOR_RECOVERY_PROBE_SECONDS` (default `5`)
- `VANTAGE_COLLECTOR_GAP_EVENT_ID_SAMPLE_SIZE` (default `32`)
- `VANTAGE_COLLECTOR_GAP_SUMMARY_INTERVAL_SECONDS` (default `60`)
- `VANTAGE_COLLECTOR_SOCKET_PATH` (default
  `/var/run/vantage-collector/collector.sock`)
- `VANTAGE_COLLECTOR_ACK_TIMEOUT_SECONDS` (default `5`)
- `VANTAGE_COLLECTOR_CONNECTION_POOL_SIZE` (default `16`)

The queue and gap aggregate are intentionally process-local and bounded. A
process exit may therefore lose queued events; the gap frame reports losses
observed while the process remains alive.

## Running the provider matrix locally

`tests/test_provider_matrix.py` is a regression matrix over the provider and
mode combinations whose usage shapes disagree with one another. Every cell is a
recorded fixture, so the matrix needs no API keys and runs on every pull
request:

```sh
pip install --editable . pytest pytest-asyncio
PYTHONPATH=. pytest tests/test_provider_matrix.py -v
```

Run a single cell while iterating on a provider:

```sh
PYTHONPATH=. pytest tests/test_provider_matrix.py -k anthropic
```

Fixtures live in `tests/fixtures/`, one JSON file per cell:

| Directory | Cells |
| --- | --- |
| `nonstreaming/` | OpenAI chat + cache, OpenAI multimodal, Anthropic post-`calculate_usage()`, Bedrock converse, Responses API + cache (with and without a `text_tokens` detail), OCR / non-token |
| `streaming/` | OpenAI SSE, Anthropic `/v1/messages` SSE, Bedrock converse, Responses API `response.completed` |

Each file carries its `usage` payload, the `expected_usage` projection, the
`expected_billable_total`, and a `note` citing the LiteLLM transform the shape
was taken from. Adding a provider means adding a JSON file — the tests
parametrize over the directory, so a new file becomes a new cell with no test
changes.

The matrix asserts three things per cell: the four projected `Usage` fields,
that those fields are mutually exclusive (they must sum to the billable total,
so no token is counted twice), and that the callback emits exactly one `started`
and one terminal event. Two further tests assert that nothing the callback
attaches to a request reaches the outbound provider body.

Two token-semantics rules are what the matrix exists to protect, both of which
have regressed before:

- **The input count is gross wherever the cache is reported nested.** LiteLLM's
  transforms fold cache reads and writes into `prompt_tokens`, and the Responses
  API counts `input_tokens_details.cached_tokens` inside `input_tokens`. Those
  buckets have to be subtracted or the cached tokens bill twice. Raw Anthropic
  usage is the exception: it puts `cache_*_input_tokens` beside a *net*
  `input_tokens`, so nothing is subtracted there.
- **`text_tokens` is not "uncached input".** It means the text modality for
  OpenAI but the net raw input for Anthropic. Reading it verbatim silently drops
  image, video, and audio input tokens, so uncached input is derived from the
  gross count instead.

For an end-to-end check against a live proxy on `:4000`, `demo/send_requests.sh`
sends tagged requests through LiteLLM; see `demo/README.md`. That path needs a
running proxy and collector and is a manual smoke test, not part of CI.
