Metadata-Version: 2.4
Name: factorial-compute
Version: 0.0.2
Summary: Typed answers from AI compute: images, masks, transcripts and text, whichever model or supplier produced them.
Project-URL: Homepage, https://factorialcompute.com
Project-URL: Repository, https://github.com/prateekvjoshi/factorial-compute
Keywords: ai,inference,api,segmentation,transcription,image-generation
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Developers
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: httpx>=0.27

# factorial-compute

One API for AI compute that returns **typed answers** — text, images, masks,
detections, depth, documents, transcripts, speech, video, 3D models, tracks —
in the same shape whichever model or supplier produced it.

```bash
pip install factorial-compute
export FACTORIAL_API_KEY=fc_...
```

```python
from factorial_compute import Client

f = Client()  # reads FACTORIAL_API_KEY

# Text, and JSON on request.
print(f.run(model="gpt-oss-120b", input="Name three uses for a forklift.").text)

# An image from a prompt.
result = f.run(model="flux-schnell", input={"prompt": "a red forklift in a warehouse"})
result.images[0].save("forklift.jpg")

# Masks: one per object, each with a box measured from the mask.
result = f.run(model="sam2-segment", input={
    "image": "warehouse.jpg",                      # a local path: uploaded for you
    "objects": [{"id": "pallet", "box": [40, 60, 300, 280]},
                {"id": "cone", "point": [512, 240]}],
})
for mask in result.masks:
    print(mask.id, mask.box)
    mask.save(f"{mask.id}.png")

# A transcript, with timestamps.
result = f.run(model="whisper", input={"audio": "meeting.mp3"})
print(result.transcript.text)
for segment in result.transcript.segments:
    print(f"{segment.start:6.1f}s  {segment.text}")
```

That is the whole surface for most work: `run()` with a model id and an
`input`, then read the typed answer.

## Rules worth knowing

- **Files go in by path, URL or bytes** under the keys `image`, `images`,
  `audio` and `video` - for vision chat models too. The client uploads them and
  sends a reference.
- **Answers are typed.** `.text`, `.images`, `.masks`, `.detections`, `.depth`,
  `.document`, `.transcript`, `.audio`, `.video`, `.mesh`, `.tracks`. Asking a
  result for the wrong kind raises, rather than returning something empty.
- **Boxes are `[x0, y0, x1, y1]` in pixels everywhere**, so a box from a vision
  model or a detector can be sent straight to a segmenter.
- **Files come back as handles.** `.save(path)` writes one; `.read()` returns
  the bytes. They stay on the server until you ask.
- **Retries are safe.** Every `run()` carries an idempotency key and is retried
  on connection failures and on 429/502/503/504, so a retry never runs — or
  bills — the work twice.
- **Errors say what to change.** A refused input raises `InvalidInput` whose
  message names the field and a value that works.
- **Usage is in the model's own unit**: tokens, `images`, `audio_seconds`,
  `characters`, `video_seconds_generated`, `object_seconds`, `meshes`.
  `result.usage.units` has it.

## Which models exist

```python
for model in f.models.catalog():
    print(model["id"], model.get("output"), model["unit"])
    print("  ", model["example"])
```

The catalogue lists the models you can call right now (`catalog(all=True)` for
everything, with `availability` saying which are offline). Each entry says what
kind of answer it returns (`output`), what every input field means (`input`),
and gives an `example` input that works as written.

## Longer work

```python
handle = f.submit(model="flux-schnell", input={"prompt": "..."})
result = handle.wait()
```

`submit` returns before the work runs and survives your process exiting.
