Metadata-Version: 2.5
Name: offpeak
Version: 0.1.2
Summary: Deadline-priced inference: give AI jobs a deadline and run them on the cheapest venue — provider batch tiers (−50%) today. Same model, same tokens, a different hour.
Project-URL: Homepage, https://github.com/offpeak-ai/offpeak
Project-URL: Repository, https://github.com/offpeak-ai/offpeak
Project-URL: Issues, https://github.com/offpeak-ai/offpeak/issues
Project-URL: Changelog, https://github.com/offpeak-ai/offpeak/releases
Author: Offpeak
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: anthropic,batch,batch-api,cost-optimization,deadline,finops,inference,llm,openai,scheduling
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: System :: Distributed Computing
Requires-Python: >=3.10
Provides-Extra: all
Requires-Dist: anthropic>=0.40; extra == 'all'
Requires-Dist: openai>=1.50; extra == 'all'
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.40; extra == 'anthropic'
Provides-Extra: dev
Requires-Dist: build; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: ruff>=0.6; extra == 'dev'
Requires-Dist: twine; extra == 'dev'
Provides-Extra: openai
Requires-Dist: openai>=1.50; extra == 'openai'
Description-Content-Type: text/markdown

# offpeak

**Deadline-priced inference.** Same model, same tokens, a different hour — for half the price.

[![CI](https://github.com/offpeak-ai/offpeak/actions/workflows/ci.yml/badge.svg)](https://github.com/offpeak-ai/offpeak/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/offpeak)](https://pypi.org/project/offpeak/)
[![License](https://img.shields.io/badge/license-Apache--2.0-blue)](LICENSE)

OpenAI, Anthropic, and Google all sell batch inference at **50% off list price**. Almost nobody uses it, because no API lets work say it can wait: every token runs "now" by default, and the batch workflow — build a file, upload, poll, download, match results back up — is enough friction that urgency gets bought by accident.

`offpeak` gives your code one new argument.

```python
import offpeak

jobs = [offpeak.job("claude-haiku-4-5", f"Summarize:\n\n{doc}") for doc in docs]

results = offpeak.run(jobs, deadline="06:00")   # done by 6am, at batch prices

print(offpeak.receipt(results))
```

```
OFFPEAK SETTLEMENT ────────────────────────────
jobs      1,000 (1,000 ok, 2 sync fallback)
sla       1,000/1,000 met
venues    anthropic:batch 1,000
tokens    12,410,332 in · 3,104,551 out
list      $27.93
paid      $14.02
captured  $13.91 (49.8%)
prices    snapshot 2026-08 — override via offpeak.prices
───────────────────────────────────────────────
```

## What it does

- **One argument, not a workflow.** `run(jobs, deadline=...)` handles batching, submission, polling, collection, and result matching across providers.
- **Deadlines are guarded, not hoped for.** If a batch hasn't landed by the time the remaining window shrinks to a risk buffer, `offpeak` cancels and re-runs the stragglers synchronously at list price. You state the deadline; it gets met.
- **Every run settles a receipt.** List cost, paid cost, captured spread — arithmetic against public price sheets, not estimates.
- **Your keys, your perimeter.** `offpeak` talks directly to the providers with your own API keys. There is no proxy and no third party in the data path.
- **Zero-dependency core.** Provider SDKs load only via extras.

## Install

```bash
pip install "offpeak[all]"        # OpenAI + Anthropic venues
pip install "offpeak[anthropic]"  # or one provider
pip install "offpeak[openai]"
```

Venues use the standard environment variables (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`), or pass a configured client: `OpenAIBatch(client=my_client)`.

## Deadlines

Deadlines are how software says "this can wait" — the full semantics live in [SPEC.md](SPEC.md).

| Form | Meaning |
| --- | --- |
| `"06:00"` | the next 6am, local time (the canonical overnight form) |
| `"4h"`, `"90m"`, `"2d"` | relative to now |
| `"2026-08-21T06:00:00-07:00"` | ISO 8601, absolute |
| `datetime` / `timedelta` / seconds | native Python forms |

## How a run works

1. Jobs are grouped by venue (`claude-*` → Anthropic Message Batches, `gpt-*`/`o*` → OpenAI Batch) and submitted at the batch tier — 50% of list.
2. `offpeak` polls the venues, backing off while the window is long.
3. When remaining time reaches the **risk buffer** (default: 15% of the window, clamped to 1–10 minutes), unfinished jobs are cancelled and re-run synchronously so the deadline holds. Set `fallback="none"` to report them instead.
4. Results come back in input order, each with a per-job `Receipt`; `offpeak.receipt(results)` settles the run.

```python
results = offpeak.run(
    jobs,
    deadline="06:00",
    fallback="sync",       # meet the deadline at list price if the batch is at risk
    risk_buffer=600,       # seconds held in reserve (optional)
)
```

## Receipts and prices

Receipts are computed against a bundled snapshot of public list prices (batch = 50% of list, as published). Providers change prices — verify and override at runtime:

```python
import offpeak

offpeak.prices.register_price("my-fine-tune", input_per_m=4.0, output_per_m=16.0)
```

Unknown models settle with `cost = None` rather than a guess.

## What this is (and the roadmap)

`offpeak` is the open client and spec for a simple claim: **intelligence has a time value**. A large share of AI work — embeddings, evals, backfills, report generation, overnight agents — has no human waiting on it, and the venues already price that patience at −50%. This library is the missing workflow.

The roadmap follows the same interface upward: more venues (Google batch, spot capacity, off-peak windows on your own GPUs), queue-latency forecasting instead of a fixed risk buffer, portfolio placement across venues, energy- and carbon-aware scheduling with per-job receipts. The venue interface (`offpeak.Venue`) is deliberately the extension point — a venue is anywhere deferred work can run.

A hosted desk that does the forecasting, cross-venue portfolio scheduling, and SLA insurance at fleet scale — payloads never leaving your perimeter — is being built by the same team. The SDK and the deadline spec stay open, Apache-2.0.

## Contributing

Issues and PRs welcome — see [CONTRIBUTING.md](CONTRIBUTING.md). Spec changes start as issues against [SPEC.md](SPEC.md).

## License

Apache-2.0 © Offpeak
