Metadata-Version: 2.4
Name: factorial-compute
Version: 0.1.0
Summary: Typed answers from AI compute: images, masks, transcripts and text, whichever model or supplier produced them.
Project-URL: Homepage, https://factorialcompute.com
Project-URL: Repository, https://github.com/prateekvjoshi/factorial-compute
Keywords: ai,inference,api,segmentation,transcription,image-generation
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Developers
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: httpx>=0.27
Requires-Dist: pillow>=10.1
Provides-Extra: images

# factorial-compute

One API for AI compute that returns **typed answers** — text, images, masks,
detections, depth, documents, transcripts, speech, video, 3D models, tracks —
in the same shape whichever model or supplier produced it.

```bash
curl -LsSf https://api.factorialcompute.com/install.sh | sh -s -- inv_...   # once per machine
pip install factorial-compute                                               # in your project
```

```python
from factorial_compute import Client

f = Client()  # the key `factorial login` saved, or FACTORIAL_API_KEY if set

# Text, and JSON on request.
print(f.run(model="gpt-oss-120b", input="Name three uses for a forklift.").text)

# An image from a prompt.
result = f.run(model="flux-schnell", input={"prompt": "a red forklift in a warehouse"})
result.images[0].save("forklift.jpg")

# Masks: one per object, each with a box measured from the mask.
result = f.run(model="sam2-segment", input={
    "image": "warehouse.jpg",                      # a local path: uploaded for you
    "objects": [{"id": "pallet", "box": [40, 60, 300, 280]},
                {"id": "cone", "point": [512, 240]}],
})
for mask in result.masks:
    print(mask.id, mask.box)
    mask.save(f"{mask.id}.png")

# A transcript, with timestamps.
result = f.run(model="whisper", input={"audio": "meeting.mp3"})
print(result.transcript.text)
for segment in result.transcript.segments:
    print(f"{segment.start:6.1f}s  {segment.text}")
```

That is the whole surface for most work: `run()` with a model id and an
`input`, then read the typed answer.

## In a coding agent

The install line sets up Claude Code: a skill that tells it what Factorial
does, and `factorial mcp` - five tools (`find_models`, `describe_model`,
`quote`, `run`, `get_result`) it can call without writing code, running as its
own key. `factorial install` sets it up again later.

## Rules worth knowing

- **Files go in by path, URL or bytes** under the keys `image`, `images`,
  `audio` and `video` - for vision chat models too. The client uploads them and
  sends a reference.
- **Answers are typed.** `.text`, `.images`, `.masks`, `.detections`, `.depth`,
  `.document`, `.transcript`, `.audio`, `.video`, `.mesh`, `.tracks`. Asking a
  result for the wrong kind raises, rather than returning something empty.
- **Boxes are `[x0, y0, x1, y1]` in pixels everywhere**, so a box from a vision
  model or a detector can be sent straight to a segmenter.
- **Files come back as handles.** `.save(path)` writes one; `.read()` returns
  the bytes. They stay on the server until you ask.
- **Retries are safe.** Every `run()` carries an idempotency key and is retried
  on connection failures and on 429/502/503/504, so a retry never runs — or
  bills — the work twice.
- **Errors say what to change.** A refused input raises `InvalidInput` whose
  message names the field and a value that works.
- **Reasoning models think before they answer**, and the thinking counts
  against `max_tokens`. Too low a cap is spent thinking: `.text` then raises
  `IncompleteAnswer` rather than returning `""`. Give reasoning models a few
  thousand tokens, or send `reasoning_effort="low"` to think less.
- **Every result says what it cost**: `result.cost.usd`, alongside
  `result.timing.duration_ms`. `f.quote(model=..., input=...)` says it before
  anything runs; `max_cost_usd=` on `run()` or `submit()` refuses a call that
  would cost more, raising `CostCapExceeded`.
- **Answers chain.** Pass an earlier answer's file as an input -
  `{"image": result.images[0]}` - and it goes by reference, never downloaded
  and uploaded again.
- **Several models as one workload**: `with f.workflow() as w:` - each
  `w.run(...)` returns at once, an input can name an earlier step's answer
  (`photo.image`, `seen.text`), and the server runs steps as their inputs
  become ready, in parallel, with one status, one cost and the whole graph.
- **Submitted work can call back**: `submit(..., webhook_url=...)` posts the
  id and status when it finishes; fetch the execution for the answer.
- **Masks and boxes draw onto the photo**: `result.masks.draw(photo, "out.png")`
  tints and numbers every mask, `mask.layer(photo, colour)` gives one
  see-through layer for stacking in HTML, `result.detections.draw(...)` boxes.
- **Usage is in the model's own unit**: tokens, `images`, `audio_seconds`,
  `characters`, `video_seconds_generated`, `object_seconds`, `meshes`.
  `result.usage.units` has it.

## Which models exist

```python
for model in f.models.catalog():
    print(model["id"], model.get("output"), model["unit"])
    print("  ", model["example"])
```

The catalogue lists the models you can call right now (`catalog(all=True)` for
everything, with `availability` saying which are offline). Each entry says what
kind of answer it returns (`output`), what every input field means (`input`),
and gives an `example` input that works as written.

## Longer work

```python
handle = f.submit(model="flux-schnell", input={"prompt": "..."})
result = handle.wait()
```

`submit` returns before the work runs and survives your process exiting.
