20 model calls. One job. One cost.

Know what your AI work costs.Without keeping what it said.

See what a real AI job costs in seconds, or run Inferrail locally.

  • No account
  • Demo mode needs no API key
  • Open source, Apache 2.0

Support ticket #42work_id: support-ticket-42

Example
$0.0148cost of this job
18model calls
8,241tokens
23stotal time
Some of the 18 calls in this example job
CallStepTokensCost
1read ticket612$0.0009
2classify380$0.0005
3look up account1,204$0.0021
4–1714 more calls5,145$0.0097
18final reply900$0.0016

An example job, for illustration. Run the live demo for a real result from a real Inferrail gateway.

What Inferrail does

It sits between your app and your model provider. Every call gets a receipt. Receipts from the same job add up to one cost.

Know what work costs

From single model calls to one workflow, customer, or job.

No prompt storage

Inferrail's receipts have no field for prompts or responses.

Set budgets

A daily cap refuses a request before the provider is called.

Works with what you use

OpenAI SDK, Anthropic SDK, LangChain, and curl.

How it works

One job, like answering a support ticket, can take many model calls. Inferrail adds them up for you.

Point your app at Inferrail

Change one base_url. Requests still go to OpenAI or Anthropic, and each call gets a receipt.

Say which job a call is for

Add one header to each call. Any name works: a ticket, an invoice, a customer.

X-Inferrail-Attribute-Work-Id: support-ticket-42

See what the job cost

Calls with the same job add up to one cost. Record how the job ended to see cost per successful outcome.

inferrail work support-ticket-42

Try it locally 30 seconds

Runs on your machine. No account. The demo needs no API key.

  1. 1 Install

    pip install inferrail
  2. 2 Run the demo

    inferrail demo
  3. 3 See your costs

    The demo prints cost per job, per customer, and per outcome.

From inferrail demo, some columns hidden

WORK ID          OUTCOME     RECEIPTS  KNOWN COST
work-contract-1  resolved    2         $0.000483
work-pending-1   undeclared  1         $0.0001
work-support-1   failed      2         $0.007825

Demo prices are made-up round numbers, labelled DEMO. No provider is called.

Connect a real appRun the gateway, point your SDK at it, tag work, set a budget.

Start the gateway

Prints the one-line base_url change for your SDK.

inferrail serve --quickstart

Point your SDK at it

OpenAI SDK: base_url="http://127.0.0.1:8000/v1". Anthropic SDK:

export ANTHROPIC_BASE_URL=http://127.0.0.1:8000

Tag calls with a job

Send this header with each call that belongs to the job.

X-Inferrail-Attribute-Work-Id: support-ticket-42

Record how the job ended

Then inferrail work support-ticket-42 shows cost and outcome together.

inferrail work outcome support-ticket-42 --status resolved

See cost by customer

Or by model, route, or any tag you attach.

inferrail report --by customer

Set a daily budget

A request that would go over is refused with HTTP 402 before the provider is called. Scoped budgets

inferrail serve --quickstart --daily-budget-usd 5.00

Open the local dashboard

Live feed, work, budgets, and copy-paste snippets. Served by the gateway itself.

inferrail serve --quickstart --app-mode

Check the privacy claim yourself

Prints the receipt schema your install is running. A schema check, not a security audit.

inferrail verify-payload-free

More: Anthropic passthrough design · exact current scope · architecture

What Inferrail keeps

Enough to tell you what the work cost. Not what was said.

Kept in Inferrail's records

  • Provider and model
  • Token counts
  • Cost, or "unknown" when it can't be priced
  • Timing and status
  • Tags you add, like a job or a customer, stored exactly as sent

Never kept

  • Your prompts
  • Model responses

The receipt schema has no field for them, and the gateway never hands message bodies to the receipt builder. Keep secrets and message text out of your tags too.

Your provider still sees the prompt

Inferrail passes your request to OpenAI or Anthropic as usual. The promise is narrower: Inferrail's own receipts and records do not contain it.

How privacy works

A receipt is a small, explicit set of fields: route, provider, model, token counts, price snapshot, cost, status, timing, and your own tags. Canary tests in this project's CI check success, error, and streaming requests. inferrail verify-payload-free prints the schema your install is running. It is a schema check, not a security audit: it can't inspect tag values you store, your logs, or your provider.

The self-hosted gateway writes records on your machine only. An optional usage ping reports lifecycle events like "installed" and "started", never anything about your traffic, and sends nothing until an endpoint is configured. Usage ping details

The self-hosted gateway reads your provider key from its own environment. The hosted trial is different: if you add a real key there, the hosted process holds it in memory, for at most 4 hours, and handles your traffic. How the trial works

What's ready today

Not everything is equally mature. Inferrail is a developer preview, so interfaces may still change before 1.0.

Available

Works today in the open-source package.

Preview

Working, and still being shaped with early users.

Labs

Experiments. Testnet only.

  • OpenAI-compatible gatewayStreaming and tool calls pass through.Available
  • Anthropic-compatible gatewayNative /v1/messages for the Anthropic SDK. Claude Code is not supported yet.Available
  • Cost receipts per callToken counts and cost. No prompt or response.Available
  • Cost per jobGroup calls by job and record outcomes.Available
  • BudgetsDaily and scoped caps that block or warn.Available
  • Hosted trialA temporary personal gateway, deleted within 24 hours.Preview
  • Local dashboardLive feed, work, budgets, snippets.Preview
  • Cost per taskSteps inside a job, with a Python helper.Preview
  • MCP serverAgents can read spend and health.Preview
  • AP invoice exception recoveryA separate capability on the same receipts.Preview
  • Work Economics and Economic AuthorityAgent economics over x402.Labs

More from Inferrail

Built on the same receipts. Separate from the core product.

Preview

AP invoice exception recovery

For invoice-extraction teams: decide whether one failed extraction gets one machine retry or goes to human review, then record what it cost.

Explore
Preview

MCP server

Two read-only tools, get_spend and get_health. An agent can check your local spend without making a model call.

See the MCP server
Labs

Inferrail Labs

Experiments in economics for autonomous agents: pay-per-call cost analysis and shared spending boundaries, on testnet.

Explore Labs

Company

A small founding team, building Inferrail in the open.

Founding Team

  • Daniel O. · Founder

    Security & AI infrastructure

    Microsoft Security Brown University

  • David O. · Head of Growth

    Williams College

Build with us

We're a small founding team building infrastructure for understanding the economics of AI work. We're especially interested in people who can help put Inferrail in the hands of developers and teams who need it.

Founding Growth & Distribution

What this role involves
  • Find and recruit early users
  • Build relationships with developer communities
  • Run creative distribution experiments
  • Turn developer interest into product adoption
  • Identify high-value design partners
  • Help tell the Inferrail story clearly
  • Learn directly from users and bring what you learn back into the product
Tell us about yourself

Your submission is sent securely through Formspree to the Inferrail team and used only to review and respond to your interest.

Investors

We're open to conversations with investors who believe in infrastructure for accountable AI economics.

Talk with Daniel

daniel_omondi@alumni.brown.edu