Metadata-Version: 2.5
Name: infimal
Version: 0.2.1
Summary: The infimal SDK, the in-container harness, and the module behind the native CLI's apps plan, apply and deploy.
License: Apache-2.0
Requires-Python: >=3.11
Requires-Dist: httpx>=0.27
Requires-Dist: pydantic>=2.7
Provides-Extra: dev
Requires-Dist: pytest-asyncio>=0.23; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Description-Content-Type: text/markdown

# infimal

The infimal Python SDK and in-container harness.

```python
from infimal import Client

client = Client.from_env()        # INFIMAL_API_KEY; INFIMAL_ENDPOINT is optional
print(client.model("deepseek/deepseek-v4-flash").chat("Explain this diff.").max_tokens(256).run().text)
```

`client.model(id)` is a request handle: `.chat(...)`, `.embeddings(...)`, `.speech(...)`,
`.images(...)` and `.video(...)` build a request, `run()` or `stream()` sends it once, and
`.capabilities()` says what the listing accepts and whether it is serving. Images and video are
jobs: `submit()` returns a `Job` you can `wait()`, `refresh()` or `cancel()`, and `client.job(id)`
adopts one after a restart.

```python
from infimal import App, Latency, Scale

app = App("voice", perf=Latency(p95_ms=300), scale=Scale(max_replicas=4))

@app.setup
def load():
    # Runs once, on a builder GPU. Weights, CUDA context, warm caches, CUDA graphs — everything
    # expensive belongs here, because what happens next is that the whole live process is frozen.
    return load_model()

@app.handler
def transcribe(model, request):
    return model(request["audio"])
```

## The environment

| Variable | Meaning |
|---|---|
| `INFIMAL_API_KEY` | the API key; required |
| `INFIMAL_ENDPOINT` | the API origin; defaults to `https://api.infimal.ai`; the host from before the rename remains supported |

`GLOW_API_KEY` and `GLOW_ENDPOINT`, the names before the rename, are still read **in 0.2 only**,
when the new ones are unset, and each read raises a one-line `DeprecationWarning` naming the
variable to set instead. 0.3 stops reading them. `GlowError` and `GlowUnreachable` resolve to
`InfimalError` and `InfimalUnreachable` for the same one release.

## The CLI

The customer CLI is the native `infimal` binary (`curl -fsSL https://infimal.ai/install.sh | sh`;
see `crates/infimal/README.md`), and it is the only one. This package installs **no console script
and has no verbs**; for the three commands that execute your module, the binary runs this package
(`python -m infimal._synth`, a private JSON-in, JSON-out subprocess described in `docs/sdk.md`), so
install the SDK where your module's dependencies are (`pip install infimal`), or point `INFIMAL_PYTHON`
at that interpreter.

```
infimal apps plan [module]          # execute your module and show what deploying would change
infimal apps apply [module]         # apply it: rebuild, or reconfigure a live placement in place
infimal apps deploy [module]        # build a snapshot and place it
infimal apps snapshots <app>        # published artifacts and how fast they restore
infimal apps status <app>           # where the endpoint is on the residency gradient right now
infimal apps logs <app>             # build logs

infimal models list                 # the price list: what each model costs you per million tokens
infimal billing balance             # credit available, credit held against in-flight requests
infimal billing fund 25 --wait      # buy credit; --wait blocks until the payment has landed
infimal usage show --since 30d [--by model]
infimal keys create|list|revoke

infimal jobs submit --surface images|video --model M --preset I1 --wait
infimal jobs get <id>               # state, measured timings, presigned artifact URLs
infimal jobs list
infimal jobs cancel <id>            # only before it starts; a running generation has been paid for
infimal jobs download <id> <dir>
```

Images and video are **jobs**, not requests: a clip is minutes of GPU time, so the submit reserves
the credit, answers `202` with an id, and you poll. `--wait` is that loop; the exit code is 1 if the
job failed.

A key is minted `full` — the whole account — unless you ask for the narrow role:

```
infimal keys create client-app --role inference   # the OpenAI surface, the model list, your usage
```

Administrators have a second set, which needs a tenant on the operator's admin list. There is no
administrative key: administration belongs to the account, so no key can grant it or carry it.

```
infimal admin models list
infimal admin models set-price <id> --input 0.15 --output 0.60 [--cache-read 0.015]
infimal admin models set-price <id> --flat 0.04 | --unpriced [--unpin]
infimal admin sources add --model <id> --name <name> --base-url <url> \
                          --upstream-model <id> --key-env PROVIDER_KEY \
                          --cost-input 0.10 --cost-output 0.30 [--wire-config @wire.json]
infimal admin sources list|rekey|delete
infimal admin grants set <tenant> <model> --granted true|false
infimal admin tenants show <tenant>
infimal admin tenants set-discount <tenant> 2000        # basis points off the sale price
infimal admin topup <tenant> 25 --reference <ref>
infimal admin audit
```

**Every command takes `--json`**, and that is the output a script should hold on to: the human text
is for a human and is allowed to change. **Every mutation takes `--plan`**, a local, redacted preview
of the request it would send. Two habits the admin commands enforce — an upstream credential is read
from an environment variable named by `--key-env`, never from an argument that would sit in the
process table and the shell history; and `--reference` on a top-up is the idempotency key, so
re-running it credits once rather than twice.

## Two things this package is opinionated about

**You never name a GPU.** There is no `gpu=` parameter. You declare `perf=Latency(p95_ms=300)` or
`perf=Throughput(rps=50)`, or just a `tier`, and the platform picks the silicon and reports back
which one it used. Pricing follows the same rule: `infimal models list` quotes per token, never
per GPU-hour, and it is the same number you are billed — whether the tokens came off our own GPUs
or were bought from a provider.

(There used to be an `estimate` verb that quoted a monthly band from a declared traffic envelope. It
priced GPU-hours against a rate card nothing charged against, so the band could never be checked
against a bill, and it is gone rather than left returning a guess.)

**`@app.setup` is a snapshot boundary, not just an init hook.** After it runs, the process — VRAM
included — is checkpointed, and every later request starts from that image instead of an import. The
corollary is the one rule it imposes: nothing that cannot be checkpointed may survive `setup`. No
open sockets, no pipes. The SDK checks the obvious cases and tells you in a sentence, because the
alternative is an opaque CRIU failure minutes into a build.

## Development

```
uv run --extra dev pytest
uv run --extra dev ruff check .
```

The ring tests also check this implementation against the Rust gateway's, byte for byte, whenever
`cargo` is available: both map the same file, and a layout drift would silently corrupt every
request.
