Metadata-Version: 2.5
Name: idiotproof
Version: 0.5.0
Summary: Software for agents that create and edit video.
Project-URL: Repository, https://github.com/4014-Labs/idiotproof
Project-URL: Issues, https://github.com/4014-Labs/idiotproof/issues
Author: 4014 Labs
License: Proprietary
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: httpx<1,>=0.27
Description-Content-Type: text/markdown

# idiotproof

Software for agents that create and edit video.

`idiotproof` is the IDEA ADK. ADK means **Agent Development Kit**: software
intended to be written and run by agents. It is designed for one-off Python scripts, CLI calls, and
headless agent workflows alongside FFmpeg. It is not an application framework
or a stable dependency for persistent software. Do not build a long-lived
service around its current interfaces; the package and dynamic tool schemas
will change as the product develops.

Bobbie plus FFmpeg is intended to be a complete headless media toolchain. Use
FFmpeg for cutting, joining, transcoding, and audio. Use Bobbie for GPU-native
vision and rendering.

## Install and authenticate

```bash
pip install idiotproof
export IDEA_SECRET_KEY=idea_sk_test_...
```

The alpha client defaults to `https://dev.67idea.com/v1`. Override it with
`IDEA_API_BASE_URL` or `Idea(base_url=...)`. A secret key belongs to one IDEA
project and environment. The API mints short-lived, workspace-scoped JWTs
internally; your code never creates or forwards those JWTs.

## Required visual-review workflow

`media_publish()` returns signed URLs specifically so an agent can inspect its
work without downloading it. Signed URLs are transient credentials, not values
that must be ignored.

After publishing media, an agent should:

1. Keep the signed URL only in the local viewer process's memory.
2. Generate the viewer response dynamically. A page may contain several URLs
   to compare multiple edits, but no signed URL belongs in an on-disk HTML file
   or manifest.
3. Serve only that directory with a server bound to `127.0.0.1`, never
   `0.0.0.0`.
4. Open the localhost URL in the available browser. If browser control is
   unavailable, still create the viewer and provide its localhost URL.
5. View the remote media through the page; do not download it merely to
   preview it.
6. Do not print, paste into chat, persist, or commit the signed URL. Stop the
   viewer process after review or expiration.

Downloading is appropriate when saving an edit for continued local work or
durable storage. It is not required for visual review.

## Python

The public vocabulary is deliberately small: one `Idea` client, opaque
`ChatId` values, dynamic runtime tools, and structured `IdeaError` failures.

```python
import asyncio
import json
from pathlib import Path

from idiotproof import Idea


async def main() -> None:
    async with Idea() as idea:
        chat_id = await idea.create_chat()
        upload = await idea.upload_file(chat_id, "clip.mp4")

        tools = await idea.get_tools()
        print([definition["id"] for definition in tools])

        # Functions and documentation come from the live MCP catalog.
        print(tools.submit_bobbie_job.__doc__)
        render_input = json.loads(Path("render.json").read_text())
        render_input.setdefault("input_media", upload["workspace_uri"])
        result = await tools.submit_bobbie_job(chat_id, **render_input)

        published = await idea.media_publish(chat_id, result["output_media"])
        # Give display_url directly to an in-memory viewer bound to 127.0.0.1.
        # That viewer is agent workflow code, not an SDK side effect. Do not
        # print the URL or write it into HTML, JSON, logs, or a manifest.


asyncio.run(main())
```

`create_chat()` returns a random opaque `ChatId` locally. The API lazily
materializes its project-scoped server workspace on the first upload or tool
call. Save the ID if another process must continue in the same workspace. Do
not put email addresses or other personal data in it.

`upload_file(chat_id, path)` probes video before network activity. An image or
video of at most 600 decoded frames follows the ordinary authorized TUS path
and returns one upload JSON object with `workspace_uri`. A longer video
automatically follows the ordered TUS metadata workflow described below and
returns an `upload_asset` JSON object with ordered `parts` and
`workspace_uris`; it never silently uploads a truncated single result.

For long videos or concurrent ordered uploads, use `upload_batch()` and
describe logical assets rather than a global upload queue:

```python
uploads = await idea.upload_batch(
    chat_id,
    [
        "long-video-a.mp4",
        ["video-b-001.mp4", "video-b-002.mp4"],
        "reference.png",
    ],
    concurrency=4,
)
```

Each top-level item is one logical asset. A bare path is one source asset; an
ordered path sequence supplies already separated pieces of one asset. Before
any request, the helper probes videos with `ffprobe` and uses FFmpeg to
re-encode every source longer than 600 frames into independently decodable
parts of at most 600 frames. It verifies that the parts preserve the complete
decoded frame count.

The helper flattens those parts in caller order and adds `uploadBatchId`,
`uploadBatchIndex`, and `uploadBatchSize` to each authenticated TUS creation.
The first creation reserves the complete contiguous block of `upload_NNN`
names; transfers may then proceed concurrently in any order. No reservation
endpoint or client-side completion throttle is involved. Results retain the
same nested asset/part order, and the helper verifies that returned basenames
form one contiguous block in that order. If an upload fails, it lets every
started operation settle and raises `IdeaBatchError`; its flat `items` retain
successes and failures in original asset/part traversal order.

`media_publish(chat_id, media_path)` calls the stable API-key-authenticated
media endpoint and returns short-lived viewing and download URLs. It is outside
the dynamic MCP catalog and does not depend on tool discovery.

`download_media(chat_id, media_path, destination)` is the durable counterpart.
It publishes internally, keeps the signed download URL in memory, verifies the
download, and atomically replaces the destination only after success. Small
objects and origins without safe range support use one streamed request. Large
objects use bounded parallel byte ranges only when the origin supplies a
content length, byte-range support, and a strong ETag. Range responses must
match their requested offsets and object identity; otherwise the helper safely
falls back or fails. `resume=True` verifies completed temporary ranges after
republishing the stable workspace path. Resume files never contain signed
URLs. In `download_and_concat()`, one shared transfer limit bounds publishing,
whole-object requests, and nested range requests so concurrency does not
multiply by the number of files.

`get_tools()` fetches each dynamic tool's name, description, and JSON Schema at
runtime. It returns an iterable catalog and generates documented async methods
such as `tools.submit_bobbie_job(...)`. Punctuation becomes `_`; use
`await tools.call(exact_name, chat_id, input_dict)` for unusual or colliding
names. Tool methods return the tool output rather than a transport wrapper.

Every discovered tool is a callable `RuntimeTool`. Call one normally, bind
shared arguments once with `partial()`, or map it over varying inputs with
bounded concurrency:

```python
render = tools.submit_bobbie_job.partial(
    chat_id,
    passes=effect,
    timeout_seconds=300,
)
jobs = await render.map(
    [
        {"input_media": "workspace:segment-001.mp4", "output_media": "one.mp4"},
        {"input_media": "workspace:segment-002.mp4", "output_media": "two.mp4"},
    ],
    concurrency=4,
)
```

Results retain input order even when calls finish out of order. The entire
input iterable is validated before any request starts, and an input key may
not duplicate a bound/common key. `tools.map(exact_name, chat_id, inputs)` is
the escape hatch for exact MCP names. `errors="collect"` returns ordered
`MapItem` values; the default waits for every item to settle and then raises
`IdeaBatchError`, whose `items` preserve both successes and failures. Mapped
tool calls are not automatically retried because they may have side effects.

Cancelling `map()` prevents queued calls from starting and cancels local waits
for calls already in flight. It cannot retract a remote operation that the
service already accepted.

Use `numbered_media()` when a pipeline needs predictable ordered workspace or
output names:

```python
from idiotproof import numbered_media

media = numbered_media(
    "workspace:video_a_rendered_{index:03d}.mp4",
    count=3,
)
# workspace:video_a_rendered_001.mp4, ...002.mp4, ...003.mp4
```

`count` must be non-negative, and the formatted values must be unique. Use a
stable sequence prefix when several source videos share a workspace. Generate
names before launching concurrent work and retain input order when collecting
results; completion time must never determine segment order. A workflow
manifest remains authoritative—zero-padding is convenient, not an ordering
guarantee. This helper is for predictable workspace and output names, never
`media_publish()` URLs: published URLs are signed and cannot be reconstructed
from a pattern.

Use `download_and_concat()` when ordered workspace videos are final and must
become one durable local video:

```python
from idiotproof import download_and_concat, numbered_media

combined = await download_and_concat(
    idea,
    chat_id,
    numbered_media(
        "workspace:video_a_rendered_{index:03d}.mp4",
        count=11,
    ),
    output="output/final.mp4",
    concurrency=4,
    resume=True,
)
```

The input must be an ordered sequence; sets, mappings, and bare strings are
rejected. Downloads may complete in any order, but index-derived local names
and caller order control concatenation. Every part is probed with `ffprobe`.
All selected video streams must have exactly equal width and height—neither
mode scales, crops, pads, or rotates a mismatch. The default
`concat_mode="copy"` requires compatible streams and never re-encodes.
Explicit `"encode"` mode may normalize other stream properties while
preserving the dimension invariant. Supplying `audio_source` ignores segment
audio and remuxes that local audio once with stream copying. It preserves the
complete video sequence even when that audio ends slightly earlier. FFmpeg
and ffprobe must be installed. The result is a `CombinedMedia` containing the
final probe summary and SHA-256.

This is a durable-output helper, not the visual-review path. Iterate by
publishing a short representative clip and comparing remote `display_url`
variants through the localhost viewer. Download and concatenate only after an
effect is selected.

All individual failures are `IdeaError`; inspect `code`, `status_code`,
`details`, and `request_id`. Never log the secret key or upload token. Signed
media URLs may be retained transiently for the localhost viewer, but should
not appear in terminal history, chat messages, durable logs, or version
control.

## CLI and MCP

Commands emit JSON so agents can compose them with scripts:

```bash
idea-adk create-chat
idea-adk upload-file "$CHAT_ID" clip.mp4
idea-adk get-tools
idea-adk call "$CHAT_ID" submit_bobbie_job --input @render.json
idea-adk map "$CHAT_ID" submit_bobbie_job --input @batch.json --concurrency 4
idea-adk media-publish "$CHAT_ID" workspace:result.mp4
```

`batch.json` contains an `inputs` array and an optional `common` object:

```json
{
  "common": {"passes": [{"kind": "fragment_shader", "shader_text": "..."}]},
  "inputs": [
    {"input_media": "workspace:one.mp4", "output_media": "one-out.mp4"},
    {"input_media": "workspace:two.mp4", "output_media": "two-out.mp4"}
  ]
}
```

The CLI returns `{"results": [...]}` by default. With `--errors collect`, it
returns ordered `items`, each containing its `index` and exactly one of
`result` or structured `error`.

Use Python `upload_batch()` when a source may exceed 600 frames or the ADK must
reserve and upload a whole multi-asset manifest.

`media-publish` returns its requested JSON to stdout, including signed URLs;
do not use that command in a logged shell. A localhost review process should
call `media_publish()` internally and retain the response only in memory.

Run `idea-adk mcp` as a stdio MCP adapter and let it inherit
`IDEA_SECRET_KEY`. It exposes `create_chat`, `upload_file`, and `get_tools`,
followed by functions generated from the live server catalog. Media publishing
remains a separate stable Python/CLI operation and is intentionally not added
to the dynamic MCP catalog.

This README is the default advice agents should receive. Project-specific
instructions belong in `AGENTS.md` for Codex or `CLAUDE.md` for Claude Code.

## Rules of thumb for agents

Uploads are currently normalized to exactly 600 frames at 3 megapixels,
slightly above 1080p. `upload_file()` detects a source beyond the input limit
and automatically creates independently decodable parts without dropping
source frames. Use `upload_batch()` directly for several logical assets or
when you want an explicitly nested result.

For a novel effect, first probe frame rate and time base and render one
representative preview clip targeting about 20 seconds while staying safely
below the service limit—590 frames is a useful cap when the maximum is 600.
Use an explicit range when supplied; otherwise the source midpoint is a
deterministic fallback, while a vision-capable agent may deliberately select a
high-motion or representative region. Upload that preview once and map several
effect variants over it. After choosing an effect, render the complete ordered
segments once. Direct full-source rendering remains reasonable for a known or
trivial effect.

Each TUS resource uses sequential, resumable chunks; core TUS requires ordered
offsets within that resource. For a long source, `upload_file()` delegates to
`upload_batch()`, which performs media-aware splitting before network activity,
reserves every output name in one transaction, and uploads those separate
resources concurrently. TUS chunks are transport details; they are not
independently decodable media segments. Upload completion order is never media
order.

Bobbie is a custom rendering engine that runs arbitrary GLSL. You describe the
pipeline but do not control bindings; Bobbie assigns them programmatically. A
pipeline can combine GLSL, smaller ML models, optimized CUDA kernels for
classical computer vision, and optional Bayesian priors over color and shape.
CUDA-GLSL interop keeps video on the GPU. Because Bobbie controls every tensor
dimension, supported pipelines should not run out of GPU memory.

Invalid requests should fail before rendering. The expected render-time
failures are a video timeout and `file too large`. Incompressible output such as
raw static can exceed 100 MB; the service deletes it. Elaborate pipelines can
serve either image/video editing or advanced computer-vision work.

A separate tool extracts frames for a VLM. The current model is Qwen-VL 27B;
custom VLMs are not supported. Bobbie requires grounded objects and coordinates
normalized to the half-open interval `[0, 1000)`.

Frame/VLM inspection plus Bobbie lets an agent move through any video. Use code
to crop, zoom, and rotate. Use vision tools to measure camera and object motion
and perform tracking, detection, and segmentation. Other catalog tools expose
metadata about uploaded files, processes, and jobs.

IDEA deletes media aggressively. If an edit matters, download it for continued
local editing or copy it immediately into durable storage such as S3.
