20 model calls. One job. One cost.
Know what your AI work costs.Without keeping what it said.
See what a real AI job costs in seconds, or run Inferrail locally.
Support ticket #42work_id: support-ticket-42
Example| Call | Step | Tokens | Cost |
|---|---|---|---|
| 1 | read ticket | 612 | $0.0009 |
| 2 | classify | 380 | $0.0005 |
| 3 | look up account | 1,204 | $0.0021 |
| 4–17 | 14 more calls | 5,145 | $0.0097 |
| 18 | final reply | 900 | $0.0016 |
An example job, for illustration. Run the live demo for a real result from a real Inferrail gateway.
What Inferrail does
It sits between your app and your model provider. Every call gets a receipt. Receipts from the same job add up to one cost.
Know what work costs
From single model calls to one workflow, customer, or job.
No prompt storage
Inferrail's receipts have no field for prompts or responses.
Set budgets
A daily cap refuses a request before the provider is called.
Works with what you use
OpenAI SDK, Anthropic SDK, Claude Code, LangChain, and curl.
How it works
One job, like answering a support ticket, can take many model calls. Inferrail adds them up for you.
Point your app at Inferrail
Change one base_url. Requests still go to OpenAI or Anthropic, and each call gets a receipt.
Say which job a call is for
Add one header to each call. Any name works: a ticket, an invoice, a customer.
X-Inferrail-Attribute-Work-Id: support-ticket-42See what the job cost
Calls with the same job add up to one cost. Record how the job ended to see cost per successful outcome.
inferrail work support-ticket-42Try it locally 30 seconds
Runs on your machine. No account. The demo needs no API key.
-
1 Install
pip install inferrail -
2 Run the demo
inferrail demo -
3 See your costs
The demo prints cost per job, per customer, and per outcome.
From inferrail demo, some columns hidden
WORK ID OUTCOME RECEIPTS KNOWN COST work-contract-1 resolved 2 $0.000483 work-pending-1 undeclared 1 $0.0001 work-support-1 failed 2 $0.007825
Demo prices are made-up round numbers, labelled DEMO. No provider is called.
Connect a real appRun the gateway, point your SDK at it, tag work, set a budget.
Start the gateway
Prints the one-line base_url change for your SDK.
inferrail serve --quickstartPoint your SDK at it
OpenAI SDK: base_url="http://127.0.0.1:8000/v1". Claude Code or the Anthropic SDK:
export ANTHROPIC_BASE_URL=http://127.0.0.1:8000Tag calls with a job
Send this header with each call that belongs to the job.
X-Inferrail-Attribute-Work-Id: support-ticket-42Record how the job ended
Then inferrail work support-ticket-42 shows cost and outcome together.
inferrail work outcome support-ticket-42 --status resolvedSee cost by customer
Or by model, route, or any tag you attach.
inferrail report --by customerSet a daily budget
A request that would go over is refused with HTTP 402 before the provider is called. Scoped budgets
inferrail serve --quickstart --daily-budget-usd 5.00Open the local dashboard
Live feed, work, budgets, and copy-paste snippets. Served by the gateway itself.
inferrail serve --quickstart --app-modeCheck the privacy claim yourself
Prints the receipt schema your install is running. A schema check, not a security audit.
inferrail verify-payload-freeMore: Anthropic passthrough design · exact current scope · architecture
What Inferrail keeps
Enough to tell you what the work cost. Not what was said.
Kept in Inferrail's records
- Provider and model
- Token counts
- Cost, or "unknown" when it can't be priced
- Timing and status
- Tags you add, like a job or a customer, stored exactly as sent
Never kept
- Your prompts
- Model responses
The receipt schema has no field for them, and the gateway never hands message bodies to the receipt builder. Keep secrets and message text out of your tags too.
Your provider still sees the prompt
Inferrail passes your request to OpenAI or Anthropic as usual. The promise is narrower: Inferrail's own receipts and records do not contain it.
How privacy works
A receipt is a small, explicit set of fields: route, provider, model, token counts, price snapshot, cost, status, timing, and your own tags. Canary tests in this project's CI check success, error, and streaming requests. inferrail verify-payload-free prints the schema your install is running. It is a schema check, not a security audit: it can't inspect tag values you store, your logs, or your provider.
The self-hosted gateway writes records on your machine only. An optional usage ping reports lifecycle events like "installed" and "started", never anything about your traffic, and sends nothing until an endpoint is configured. Usage ping details
The self-hosted gateway reads your provider key from its own environment. The hosted trial is different: if you add a real key there, the hosted process holds it in memory, for at most 4 hours, and handles your traffic. How the trial works
What's ready today
Not everything is equally mature. Inferrail is a developer preview, so interfaces may still change before 1.0.
Works today in the open-source package.
Working, and still being shaped with early users.
Experiments. Testnet only.
- OpenAI-compatible gatewayStreaming and tool calls pass through.Available
- Anthropic-compatible gatewayNative
/v1/messages. Works with Claude Code.Available - Cost receipts per callToken counts and cost. No prompt or response.Available
- Cost per jobGroup calls by job and record outcomes.Available
- BudgetsDaily and scoped caps that block or warn.Available
- Hosted trialA temporary personal gateway, deleted within 24 hours.Preview
- Local dashboardLive feed, work, budgets, snippets.Preview
- Cost per taskSteps inside a job, with a Python helper.Preview
- MCP serverAgents can read spend and health.Preview
- AP invoice exception recoveryA separate capability on the same receipts.Preview
- Work Economics and Economic AuthorityAgent economics over x402.Labs
More from Inferrail
Built on the same receipts. Separate from the core product.
AP invoice exception recovery
For invoice-extraction teams: decide whether one failed extraction gets one machine retry or goes to human review, then record what it cost.
ExploreMCP server
Two read-only tools, get_spend and get_health. An agent can check your local spend without making a model call.
Inferrail Labs
Experiments in economics for autonomous agents: pay-per-call cost analysis and shared spending boundaries, on testnet.
Explore Labs