Metadata-Version: 2.4
Name: moviestar
Version: 0.4.2
Summary: Video editing CLI for AI agents. Compose scenes, animate camera moves, retime footage, mix audio, burn captions, and render with FFmpeg.
Author: James Dillard
License: BSL-1.1
Keywords: video,ffmpeg,ai,agents,cli,transcript,whisper,editing
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Topic :: Multimedia :: Video
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: click>=8.0
Requires-Dist: faster-whisper>=1.0
Requires-Dist: rapidfuzz>=3.0
Provides-Extra: diarize
Requires-Dist: pyannote.audio>=3.0; extra == "diarize"

# moviestar

**Video editing CLI for AI agents.** Give your coding agent the ability to edit video.

Your agent loads one or many videos, fuzzy-searches the transcript, composes a timeline across multiple cameras, animates per-scene camera moves, retimes footage, mixes source audio with voiceover and music, burns in captions with spoken-word highlighting, and renders with FFmpeg — or extracts any number of standalone clips in one pass. All output is structured JSON. Source files are never modified — edits are described in a declarative spec and applied at render time.

## Quick start

Make sure FFmpeg is on PATH (`brew install ffmpeg` on Mac), create a Python 3.10+ venv, then paste this into Claude Code, Cursor, or any coding agent:

```
Create a Python 3.10+ virtual environment, then:

pip install moviestar
moviestar --help

Read the --help output — it's your operational briefing.

Then load the video at ~/Downloads/podcast.mp4. Find the moment where
the host says "the bottom line, here's what I think." Trim to just
that sentence with --snap-to-words so the cuts don't land mid-word.
Export the result to ~/Desktop/clip.mp4 and tell me how long it is.
```

The agent will work through `load → find → trim → export`, returning structured JSON at every step. No GUI, no timeline, no manual scrubbing.

`python -m moviestar` is an equivalent entry point when an agent wants to
guarantee it is invoking MovieStar from the active Python environment.

For a multi-camera recording, the agent can `load cam1.mp4 cam2.mp4 screenshare.mp4`, `concat` ranges from each into one timeline, pick whose mic plays with `--audio-from`, choose a global layout with `--layout`, or use `scenes set` for layout changes over time. Then `captions generate` burns word-highlighted captions from the transcript into the final render.

For a product demo, `scenes motion` adds source-relative zooms and pans plus speed, target-duration, and hold pacing to individual scene slots. `audio` places voiceover and music on the finished timeline, adjusts gain and fades, loops beds, and ducks one layer under another.

## What's in MovieStar

**Browse:**
- `moviestar load` — index one or more videos, transcribe with local Whisper. `--as <name>` names sources; `--add` appends to an existing project; `--no-download` guarantees a hermetic run.
- `moviestar skim` — fast browse: thumbnails + transcript over a range (`--text-only` for transcript alone).
- `moviestar inspect` — dense thumbnails on demand at a configurable interval.
- `moviestar storyboard` — one labeled contact sheet that maps a whole video or result-time range at a glance.
- `moviestar watch` — extract an MP4 segment for multimodal model analysis.
- `moviestar status` — current project state at a glance.
- `moviestar history` — step-by-step edit lineage for a source.

**Edit (multi-source):**
- `moviestar trim` — append a trim to a source's edit spec (result-time semantics; stacks compose).
- `moviestar cut` — remove a range from the middle of a source's result.
- `moviestar concat` — stitch ranges from one or more sources into a single composition. `--audio-from <source>` routes which source's audio plays across the cut; `--canvas`, `--layout`, `--slot`, and `--framing` build canvas-aware visual compositions.
- `moviestar layouts` — inspect preset layout regions and sample images before choosing a layout.
- `moviestar scenes set` / `moviestar scenes list` — author and read ordered scene-layout compositions where each scene can use its own preset layout, slots, source ranges, framing, and audio route. For large compositions, pass editable JSON with `moviestar scenes set scenes.json --dry-run`, review the resolved plan, then re-run without `--dry-run`.
- `moviestar scenes motion dump` / `moviestar scenes motion set` — round-trip per-slot pacing and camera state as editable JSON. Speed a source-local range up or down, fit it to an exact result duration, hold a frame, or define source-relative zoom/pan targets. `--dry-run` reports calculated peer pacing, time maps, normalized/pixel targets, derived crops/zoom, and pacing-driven camera rebases. Screenshot, inspect, watch, and export all render the same resolved camera and pacing plan. Once motion exists, scene/slot IDs preserve it across explicit JSON renames; deletions report exactly which attached motion records were removed.
- `moviestar undo` — pop the last operation (or the last concat).
- `moviestar spec` — show the current spec (or `--edit` to replace, `--reset` to clear).
- `moviestar find` — fuzzy-search the transcript across every source; `--context` returns the surrounding timestamped segments.

Every source-specific command takes `--source <id>`, required only when a project has more than one source.

**Captions and text overlays:**
- `moviestar captions generate` — compile the transcript into caption overlays, following each scene's audio routing. `--highlight spoken-word` colors the word being spoken, timed from the transcript's word timing.
- `moviestar captions import` — import `.srt` / `.vtt` caption files as overlays (cue timing is result-time; line breaks preserved).
- `moviestar captions rules add --merge "ground -truthing=ground-truthing"` / `--replace "mispelled=misspelled"` — persist exact project-level token corrections that every future `captions generate` applies before cue grouping. Use `captions rules list` and `captions rules remove <rule-id>` to inspect or remove them.
- `moviestar overlays add` — one timed text overlay: title, lower third, label, or creative emphasis text. Position presets or normalized `--x/--y`, `--z-index` stacking, style presets plus a CSS-like subset (`--css "font-size: 72px; color: white; -moviestar-stroke: 4px black; transform: rotate(-8deg)"`).
- `moviestar overlays dump` / `moviestar overlays set` — round-trip the complete overlay state as editable JSON: fix caption text, shift timing, restyle, then atomically replace after validation.
- `moviestar fonts list` / `moviestar fonts add path/to/font.ttf --name "My Font"` — see bundled fonts and install custom `.ttf`, `.otf`, or `.ttc` files without touching package internals. Inside a loaded project, fonts install to `moviestar/fonts`; outside a project, they install to `~/.moviestar/fonts`. Use them from captions or overlays with `--css "font-family: My Font"`.
- Overlays render on every verification surface — `screenshot`, `inspect`, `watch`, and `export` all burn them in, with bundled or installed fonts resolved before rendering.

**Audio mix:**
- `moviestar audio source` — inspect, mute, or adjust the gain of routed source audio.
- `moviestar audio add` — place voiceover or music against result time with source trimming, gain, fades, optional looping, and signal-driven ducking.
- `moviestar audio dump` / `moviestar audio set` — round-trip the complete mix as editable JSON and validate it with `--dry-run` before rendering.
- `watch` and `export` render the same resolved mix. Non-default mixes use a final clipping-protection limiter, with optional loudness normalization applied after the complete mix.

**Output:**
- `moviestar export` — render the composition, scene composition, layout composition, or a single source's edit to MP4, `moviestar-clips/export.mp4` by default. Use `--loudness-target -14` to normalize shorts-style audio, and `--audio-join-fade 0.08` to smooth sequential cut/scene joins without changing result timing.
- `moviestar clip` — extract one standalone clip to its own MP4 (source-time, leaves the edit spec untouched).
- `moviestar batch` — extract many standalone clips from a JSON recipe in one atomic, frame-exact pass.
- `moviestar clean` — preview or delete generated media artifacts.
- `moviestar screenshot` — single frame at a timecode (project-aware: `--at` is in result-time).

**Always-available:**
- `moviestar probe` — ffprobe metadata as JSON. Add `--loudness` (optionally with `--from` / `--to`) for structured integrated loudness and peak metrics.
- `moviestar models pull <model>` — pre-download a Whisper model before a load, CI job, or offline session.
- `moviestar feedback "message"` — anonymously send product feedback to the MovieStar team. The payload contains only the message and installed MovieStar version; it never includes project files, paths, session logs, or command history.
- `--dry-run` on every expensive or high-impact command (`load`, `inspect`, `watch`, `export`, `clip`, `batch`, `concat`, `scenes set`, `spec --edit`).

Image-producing commands return files by default: read `out` or each
`thumbnails[].path` with the agent's image/file tool. Pass `--inline` only
when a custom executor converts the returned Base64 bytes into image content
blocks; ordinary CLI harnesses otherwise ingest those bytes as text.

Run `moviestar --help` for the full command list.

## Why moviestar

Video editing tools are built for humans with GUIs. Agents don't have hands on a timeline or eyes on a canvas. moviestar is the hands; the agent is the brain.

- **Verbose by default.** Every command returns rich structured JSON — agents can discard what they don't need, but can't invent data the CLI didn't provide.
- **Deterministic.** Same input + same parameters = same output. No randomness, no hidden model calls.
- **Result-time semantics.** A second trim narrows the current result, not the original source. `find` matches resolve to both source-time and result-time so screenshots and exports compose cleanly.
- **Multi-source native.** Load many cameras into one workspace, compose a timeline across them, and route audio per-composition — the same structured CLI the single-source flow uses.
- **Non-destructive.** Source files are never modified.

Full design and product principles in the project docs.

## Requirements

- **Python 3.10+**
- **FFmpeg** on `PATH` (`brew install ffmpeg` / `apt install ffmpeg`)

The package weighs ~210MB on install — `faster-whisper` ships local transcription out of the box (no API keys, no cloud round-trip). Diarization is opt-in: `pip install moviestar[diarize]`.

## Status

v0.4 adds per-scene camera motion and pacing plus result-time audio mixing to the v0.3 visual-composition arc. M20 and M21 engineering are complete, and v0.4 is ready for people to use while canonical real-QuickTime, talking-head, voiceover, and music validation continues. Post-v0.4 main adds the storyboard overview, paths-first visual envelopes, and `python -m moviestar` invocation. The roadmap of what's next lives with the project.
