INTERNAL MEMO SpacePilot Studio · Master Product Memorandum
Date August 21, 2026
Author Principal PM & Swarm Architect
Subject SpacePilot Cinema Workstation & Zero-Markup Compute Broker
Status v2.3.0 Production Release (80+ Tests Passing)

SpacePilot Studio: Executive Strategy, Zero-Markup Compute Broker & Autonomous ADLC Blueprint

"SkyPilot pilots your cloud servers. SpacePilot pilots your generative cinema." Dismantling the 20x SaaS credit-markup model via resident AWS/Shadeform Spot orchestration, zero-build FastAPI architecture, native FastMCP servers, and autonomous multi-agent swarms.

$0.012 Cost per 4s Video Take
$0.0004 Kokoro VO per Min (108× vs ElevenLabs)
0.0s Resident VRAM Cold Start
63 / 63 Verified Test Suite (100%)
Section 1.0

The Core Product Vision & 4-Layer Open-Weight Architecture

The SpacePilot Cosmic Thesis: "SkyPilot routes your cloud servers. SpacePilot pilots your generative cinema into deep space."

🌌 The 5 Cosmic Horizons of Generative Cinema

🌍 1. GROUND

Zero-cost local workstation, Apple Silicon MLX, zero-build ES6 Studio UI.

☁️ 2. SKY

Multi-cloud spot arbitrage via SkyPilot (AWS, GCP, RunPod at $0.75/hr).

🚀 3. SPACE

3D Latent camera trajectories, FLF2V morphing, 0.0s resident Float8 VRAM.

🌌 4. GALAXY

Real-time 60fps streaming diffusion + 4D Volumetric Gaussian Splats.

🪐 5. UNIVERSE

Autonomous AI agent film factories with HTTP 402 micro-USDC settlement.

A forensic landscape audit of open-source tools and adjacent platforms revealing why no single project integrates the entire 5-horizon stack:

Project / Platform Category What It Does Well (Strengths) What It Lacks vs. SpacePilot Studio Vision
SkyPilot (UC Berkeley) Multi-Cloud Orchestrator 1-click launch on AWS, RunPod, GCP; automatic spot recovery and lowest-cost routing; clean Python/CLI API. CLI-only / No Creative UI: Zero visual studio, no built-in AI coding agent, no video/audio media workflow.
dstack (dstack.ai) AI Control Plane BYOC fleet management (AWS, GCP, RunPod, K8s); dashboard for GPU clusters & volumes; serves vLLM/Ollama. Generic LLM focus: Not designed for generative media; no embedded coding agent; no creative directing surface.
BentoML / OpenLLM Model Serving & BYOC Packages Diffusers/PyTorch models into APIs; warm worker management; BentoCloud VPC deployments. Developer-only: No visual creator frontend; manual container setup; no prompt/media director capabilities.
ComfyUI + ComfyDeploy Diffusion Node Graph Unmatched granular control over diffusion pipelines; community nodes for LTX, Wan, Flux; headless API export. Spaghetti Node Graph: Complex learning curve; no cloud GPU lifecycle / spot cost management; no AI debugger.
Open WebUI + LiteLLM LLM Gateway & UI Multi-model router & dashboard; manages Ollama/vLLM + OpenAI APIs; built-in web search and tools. Chat-only: Focused on text LLMs; cannot orchestrate raw EC2/RunPod GPU spot hardware or heavy video diffusion.
Modal Labs Serverless GPU Infra Clean Python decorators for GPU/RAM; sub-second cold starts; warm container holding. Proprietary Cloud: Pay-per-second markup over raw spot compute; no open-source BYOC; no creator frontend studio.
Figure 1.1: Complete End-to-End System Design & Dataflow Architecture PLANE 1 CREATOR & WORKSTATION EDGE Directing Controls • 3D Camera Trajectory Compass • FLF2V Morph Dropzones • Aesthetic Style Matrix • Real-Time Dial Scrubber NLE & Timeline • J-K-L Frame Scrubber • Audio Sync Canvas Telemetry • Live 1s Cost Odometer • SSE Status Streaming PLANE 2 EDGE & SECURITY GATEWAY Zero-Build Control Plane • Auth / Bearer Gateway • confirm:true Safety Barrier • Path Traversal Sandbox • 50MB Binary Ingestion • SSE Telemetry Streamer Media & Audio Engine • In-Process Kokoro TTS • Audio LUFS Normalizer • Curated BGM Cache Router PLANE 3 AUTONOMOUS MLOPS MESH The Director Core • Gemini Narrative Decomposer • 6-8 Shots Script Engine • Multi-Take Seed Consistency • Continuity Locker Tracker MLOps & Safety • CUDA VRAM Live Watchdog • OOM Crash Prevention • Dead Man's Auto-Idle • 20m Cost Guard Shutdown • Spot Price Arbitrage Poller PLANE 4 ELASTIC DATA PLANE Resident GPU Worker • AWS Spot L40S/A10G • PyTorch Resident Daemon • `ltx_worker.py` (Active) • Quantized Float8 LTX-2.5 • Warm 48GB VRAM (0.0s cold) Local Storage Node • 230GB NVMe Scratch Disk • `/scratch/hf` Model Cache • `/scratch/out` Output Dir • Compressed Rsync Ready REST / SSE RPC Bearer Token HTTP:8088 (Inference Requests) Compressed Diff / Rsync Video Stream directly to workstation Telemetry Heartbeat (2s)
Section 2.0

The Economic Moat & Unit Cost Arbitrage

The Strategic Insight: Every incumbent in generative video (Runway, Pika, Kling, Luma) acts as a centralized cloud middleman, marking up GPU cycles by 500% to 2,000% through artificial credit currencies. SpacePilot Studio inverts this by offering direct, single-tenant BYOC (Bring Your Own Compute) Spot GPU execution (~$0.75/hr = $0.01/take) combined with resident model memory, zero cold starts, and in-process speech synthesis.
Figure 2.1: The Three Pillars of SpacePilot's Structural Cost Moat 01 · RAW COMPUTE AWS Spot Fleet ($0.75/h) • ~$0.01 per 4s take • 10×-20× cheaper than credits • Direct single-tenant allocation 02 · RESIDENT DAEMON Warm VRAM Model Core • LTX-2.5 warm in 48GB VRAM • 230GB NVMe local scratch • 0.0s request cold start 03 · FULL PRODUCTION Voice + BGM + Camera • Kokoro TTS ($0.0004/min) • Curated BGM router cache • Sidechain auto-ducking
The Macro Paradigm Shift: The Emergence of the "AI Inference Market"
"Interfaces and frontends are generated at the speed of thought for free. Backend glue code is free. SOTA intelligence and open weights are free. Local homelab inference is free. The only un-fakeable commodity left on earth is GPU compute. We are witnessing the transformation of AI inference into a financialized commodity market—complete with dynamic pricing, warm-capacity priority routing, spot auctions, and cross-cloud arbitrage."
Tier / Pricing Mechanism Clearing Model Creator SLA & Latency Economic Advantage vs. SaaS
🔥 On-Demand Warm (Surge) Dynamic surge pricing based on real-time cluster load. Instant / 0.0s cold start: Model held warm in 48GB VRAM. Users willingly pay $0.03 instead of $0.01 for immediate director take iteration without 5-minute boot penalties.
⚡ Spot Priority Bidded AWS/RunPod spot capacity + double-digit % orchestration fee. ~30s render: Dedicated single-tenant L40S execution. 10x–20x cheaper: Creator captures pure spot savings ($0.75/hr) rather than locked $50/mo credit bundles.
🌙 Trough Batch (Overnight) Preemptible queue absorbing idle valley capacity. Asynchronous: 60-shot storyboard film renders in background. Sub-cent floor pricing ($0.006/take), soaking up surplus compute when global demand drops.
🌐 Cross-Cloud Arbitrage Multi-provider router (AWS Spot ⇆ RunPod ⇆ Lambda ⇆ Homelab). Dynamic Failover: Auto-routes if spot capacity pre-empts. Zero vendor lock-in. Real-time arbitrage between geographical spot price drops.

2.2 The Universal AI Distribution Mesh: LiteLLM, MCP & Agent Skills

SpacePilot Studio is designed not merely as a human creator tool, but as the universal video generation runtime for the global AI agent economy. By adhering to open protocol standards (LiteLLM Proxy, Model Context Protocol, and Agent Skills), SpacePilot can be discovered, scripted, and rented by any frontier LLM (Claude, GPT-4o, Gemini 2.0 Pro), open-source router (Ollama, vLLM, DeepSeek), or coding agent (Cursor, Claude Code, Antigravity) natively across four discovery interfaces:

🤖 1. LiteLLM & OpenAI-Compatible Gateway

SpacePilot exposes a drop-in /v1/images/generations and /v1/video/generations OpenAI-compatible proxy interface. Any developer using LiteLLM, LangChain, or LlamaIndex can point their existing SDK to SpacePilot by changing only api_base, gaining instant access to quantized LTX-2.5 and Kokoro TTS fleets with zero code refactoring.

🔌 2. Native Model Context Protocol (MCP) Server

A native MCP server (spacepilot-mcp via Python FastMCP) exposes granular tools: spacepilot_probe_hardware(), spacepilot_recommend_models(), and spacepilot_measurements() directly into Cursor, Claude Desktop, and IDE sidecars without custom integrations.

🧠 3. Agent Skills Standard (`SKILL.md`)

Packages SpacePilot's procedural directing knowledge into an open Agent Skills format (agentskills.io standard). AI director agents load progressive guidance for 3D camera trajectory math, STG guidance tuning, and Kokoro voice LUFS normalization on-demand without context window exhaustion.

⚡ 4. Python SDK & Zero-Friction CLI (`pluto`)

Deterministic command-line interface for CI/CD pipelines, automated marketing bots, and programmatic video rendering: pluto create --prompt "..." --camera-pan right --duration 4.0 --json. Outputs clean JSON artifacts, MP4 byte streams, and machine-readable metadata.

2.3 The Universal Neural Fabric: Inference, LoRA Fine-Tuning, Training & Hub Ecosystem

SpacePilot is architected as a universal, lifecycle-complete neural execution fabric. Rather than building a walled garden or attempting to replicate HuggingFace, SpacePilot acts as the universal runtime player and spot orchestrator for any open-weight model in existence—bridging inference, rapid LoRA fine-tuning, full checkpoint training, and continuous hub synchronization.

⚡ 1. UNIVERSAL INFERENCE

Paste any HuggingFace model repo (hf.co/...), Civitai link, or safetensors URL. SpacePilot calculates VRAM footprints, applies Quantized Float8/Int4 quantization, and boots an optimized runtime (vLLM-Omni, SGLang, ONNX, or MLX) with 0.0s cold start.

🎨 2. 1-CLICK LoRA & FINE-TUNING

Upload 10-20 reference images or style clips directly in the Cockpit. SpacePilot spins up a cheap spot GPU ($0.40/hr), runs a parameterized LoRA fine-tuning run (PEFT / Kohya / Diffusers Trainer), and registers the hot-swappable adapter directly to your studio picker.

🏋️ 3. SPOT TRAINING & CHECKPOINTS

Declarative multi-node spot training recipes via SkyPilot. Automated NVMe checkpoint sync to Cloudflare R2 / S3, continuous loss curve telemetry streaming to the Cockpit, and automated spot preemption checkpoint recovery.

🏆 4. LEADERBOARD RECIPE GALLERY

Curated, battle-tested hardware recipes for the world's top open foundation models (LTX-2.5, Wan2.1, HunyuanVideo, DeepSeek-R1, Qwen2.5-VL, Kokoro-82M). 1-click launch with optimal spot instance types and pre-calculated hourly cost quotes.

Section 3.0

Granular Product Features & Directing Suite

P0 · Core Control
🎥 3D Camera Trajectory Compass

Visual dial controller in `/create` translating user adjustments (Pan, Tilt, Dolly Zoom, Roll, Velocity) into deterministic spatial tokens injected into LTX-2.5 guidance vectors.

Target: spacepilot/web_api.py + web/create.js
P0 · Continuity
🔄 Video Extension & Branching (+4s)

1-click extension on any completed take. Extracts the final video frame losslessly via FFmpeg and seeds it as the keyframe for the next take, maintaining temporal continuity across scenes.

Target: Frame -1 Extraction -> I2V Pipeline
P0 · Cost Guard
🛡️ Dead Man's Switch (Auto-Shutdown)

Watchdog on the GPU Cockpit API. If no jobs are active or queued for 20 minutes (configurable), automatically executes `pluto terminate` to halt AWS Spot billing.

Target: Background Daemon + Cockpit UI
P1 · Transformation
🖼️ First-Frame & Last-Frame (FLF2V)

Dual image dropzones (`[Start Keyframe]` ──> `[Motion]` ──> `[End Keyframe]`), allowing creators to direct complex physical morphs and cinematic scene-to-scene transitions.

Target: Dual Aspect Dropzones + Worker Hook
P1 · Exploration
🎲 4-Take Director Grid (2×2 Batching)

Single-click trigger generating 4 seed-varied takes of the same prompt in parallel on the L40S, displayed in a responsive gallery for rapid visual branching.

Target: Batch Size = 4 · Seed Variance Matrix
P0 · Infra & Onboarding
🌐 Unified Multi-Cloud Provider Hub & Wizard

Single-click choice between Shadeform Multi-Cloud (20+ providers auto-spot routing), AWS Direct Spot ($0.75/hr), RunPod, or Local Homelab ($0.00/hr) with real-time price & fleet matrix comparison.

Target: web/cockpit.html + web/sidebar.js
P0 · Internal Admin
🔍 Cockpit Inspect Mode (Web SSH & Diagnostics)

Full-screen diagnostic drawer featuring interactive Web SSH terminal (xterm.js + WebSocket Initial Auth Frame), live hardware stats (`nvidia-smi`, `df -h`), and 1-click worker daemon restart actions.

Target: WS /api/gpu/inspect/shell + PTY Async Bridge
P2 · Autonomous
🎭 Multi-Shot Script-to-Storyboard

Paste a full script -> Gemini Flash decomposes the narrative into 6–8 shot prompts with locked character seeds and camera motion cues for batch production.

Target: Domain Pack #3 Studio Directing
Section 4.0

39-Rival Competitive Topology & Market Matrix

Mapping the generative video ecosystem across two critical dimensions: World Simulation Fidelity (cinematic motion vs static avatars) and Production Control (casual 1-click generators vs pro workstations).

Figure 4.1: 2D Generative Video Competitive Topology ($45.8B TAM) ▲ CINEMATIC & WORLD SIMULATION ▼ AVATAR & SPOKESPERSON CONSUMER PRO WORKSTATION 🎬 PLUTO STUDIO • LTX-2.5 Resident VRAM • Voice + BGM + Spot Fleet Runway Gen-3 Kling · Luma Dream Grok Imagine Higgsfield · Pika HeyGen · Synthesia Talking Heads / LMS Diffusion Studio · Remotion
Integration Vector Supported Consumers Invocation Protocol Pricing & Settlement Rail
LiteLLM Proxy OpenAI SDK, LangChain, CrewAI, AutoGen, Vercel AI SDK REST POST /v1/video/generations API Key / Usage Balance ($0.04/take)
Model Context Protocol (MCP) Cursor, Claude Desktop, Windsurf, Antigravity, OpenCode JSON-RPC 2.0 over Stdio / SSE Session Token / Enterprise MoR
Agent Skills (`SKILL.md`) Claude Code, OpenAI Swarms, Custom LLM Subagents Progressive Disclosure YAML + Markdown Free Procedural Knowledge / Open Standard
Competitor Core Category Key Strengths Fatal Weaknesses SpacePilot's Asymmetric Wedge
Runway (Gen-3) Cloud AI Suite High brand recognition, Act-One character capture, Motion Brush. Prohibitive credit pricing ($76/mo), shared queue delays, closed model. Direct AWS Compute: 10x-20x cheaper per take with dedicated zero-wait GPU.
Grok Imagine (xAI) Social Real-Time Extreme generation speed, unmoderated physics, viral X integration. Single-shot only; zero camera dials, audio synthesis, multi-shot sequencing or project state. Full Production Suite: Multi-shot storyboard, Kokoro voice, BGM, and 4K mastering.
HeyGen / Synthesia Corporate Avatars Hyper-realistic avatars, multi-language dubbing, enterprise procurement dominance. Stiff talking-head format. Incapable of cinematic world motion, camera sweeps, or VFX. Cinematic Storytelling: Dynamic camera physics for VFX, creators, and indie studios.
Diffusion Studio Web-Native NLE Smooth WebCodecs canvas, clean modern UI, developer-first tooling. Relies on third-party API wrappers; no native resident GPU orchestration or diffusion core. Integrated Stack: Direct resident model control with live telemetry and auto-shutdown.
Remotion Code-First Video Deterministic React video rendering, perfect parameter keyframing, dev loyalty. Requires full TypeScript coding for every cut; CPU Lambda rendering cannot do live diffusion. Visual Generative Directing: Visual prompt-to-video workflow with underlying reproducibility.
Section 5.0

Fleet, Security & Infrastructure Blueprint

Figure 5.1: SpacePilot Studio Zero-Build Request & Worker Topology STUDIO CLIENT • /create (Prompt/I2V) • /cockpit (Telemetry) • window.mvDialog • SSE Log Stream REST / SSE FASTAPI BACKEND • Port 8000 (Uvicorn) • Auth & confirm: true Guard • Kokoro TTS + LUFS • 55 Unit Tests Verified HTTP:8088 AWS SPOT L40S • ltx_worker.py (Flask) • Warm VRAM (Float8) • /scratch NVMe 230GB • Rsync Output Stream
Section 6.0

4-Sprint Execution Roadmap & Delivery Gantt

Sprint Phase Strategic Objective Key Deliverables Success Metrics
Sprint 1 · ✅ Complete Core Runtime & Fleet Foundation • Dead Man's Switch (Auto-Shutdown watchdog on inactivity)
• 1-Second Live GPU Uptime & Accrued Cost Odometer
• 3D Camera Trajectory Compass & Spatial Guidance Dials
• Dynamic Cost Quoting ($0.00 Local vs ~$0.04 Spot)
• Multi-Cloud Provider Hub & 3-Card Onboarding Wizard
63/63 Full Integration Tests Passing; Zero runaway billing; Dynamic Local vs Spot quotes live.
Sprint 2 · ⚡ Active Swarms Directing Suite & Connectivity • Cockpit Inspect Mode (Interactive Web SSH & Remote Diagnostics)
• Clip Extension & Temporal Continuity (+4s Frame-1 Chaining)
• Native FastMCP Server (`spacepilot-mcp`) + Open Agent Skills (`SKILL.md`)
• First-Frame + Last-Frame (FLF2V) dual keyframe dropzones
Browser-based SSH terminal for admin box triage; seamless multi-clip video extensions; MCP agent discovery.
Sprint 3 Autonomous Director & Batch Mesh • 4-Take 2×2 Exploration Grid (Parallel L40S Batching)
• Multi-Shot Script-to-Storyboard Decomposer (Domain Pack #3)
• Kokoro voiceover studio integration in `/create`
• Sidechain BGM auto-ducking on take preview
User pastes 60s script and gets full multi-take storyboard with coherent seeds and normalized broadcast audio.
Sprint 4 Global Settlement & Monetization • Human & Enterprise Merchant of Record (Polar.sh / Dodo Payments)
• Autonomous AI Agent Micropayments via HTTP 402 (`x402` Header)
• Self-Hosted Usage & Entitlement Metering Engine (La

6.1 Autonomous Swarm High-Level Strategic Roadmap

Multi-Wave macro delivery horizons. For live interactive agent task dispatch, real-time telemetry, and PR cards, visit SpacePilot Oven 🛸.

Open Live Oven Board →
Phase / Wave Strategic Focus Core Deliverables Status & Target
Wave 1–3
Foundations
Core Cinema • Quantized LTX-2.5 0.0s resident daemon on 48GB VRAM
• Kokoro-82M broadcast voice & -16 LUFS sidechain normalization
• FLF2V dual keyframe morphing & 4-Take 2×2 director grid
• Web SSH Cockpit PTY terminal & hardware monitoring
✅ Shipped (v2.3.0)
Wave 4
Scale & DX
Enterprise Fleet • SkyPilot multi-cloud spot arbitrage (AWS, RunPod, Shadeform)
• Gemini script & storyboard decomposer (60s narrative breakdown)
• Top open model leaderboards (Wan 2.1, Hunyuan, DeepSeek)
• Polar.sh merchant of record & HTTP 402 agent micropayments
🔍 Review & Merge
Wave 5
Fine-Tuning
Creator PEFT • 1-Click LoRA Studio (Spot PEFT fine-tune for ~$0.40)
• R2 spot checkpoint streaming & preemption auto-recovery
• Fullstack refactor (FastAPI modular routing + React UI migration)
• Vector transcript memory & agentic session search
🍳 In Development
Horizon 6+
Spatial & Latent
Next-Gen • 4D Gaussian Splatting for Vision Pro / Quest spatial cinema
• 60 FPS sub-50ms StreamDiffusion WebRTC latent streaming
• Autonomous Multi-Agent Screenplay Collaboration Swarm
• Enterprise SAML/SSO & zero-knowledge encrypted storage
🧊 Backlog
Section 7.0

System Changelog & Version Ledger

Chronological audit log of all major architectural additions, safety protocols, and full-stack merges pushed to origin/main:

Version & Date Category Summary of Changes Git Commit & Test SLA
v2.3.0
Aug 21, 2026
Production Release • SpacePilot Cinema Workstation Rebranding: "SkyPilot pilots your cloud servers. SpacePilot pilots your generative cinema."
• First-Frame & Last-Frame (FLF2V) Dual Keyframe Morphing: Responsive start/end dropzones, bidirectional aspect mismatch guards, and local FFmpeg cross-dissolve video generator.
• 4-Take Director 2×2 Exploration Grid: Deterministic seed batch stepping (`req.seed + i`), synchronized 2×2 playback player, and take-group lineage.
• Cockpit Inspect Mode: Interactive Web SSH Terminal (`xterm.js`), non-blocking WebSocket PTY bridge, and remote `nvidia-smi` hardware diagnostics.
• Temporal Video Extension (+4s): Lossless Frame-1 extraction and seamless scene continuation with path-traversal sandboxing.
• Native FastMCP Server & Open Agent Skills: FastMCP tool suite (`spacepilot/mcp_server.py`) and `SKILL.md` director standard for Claude, Cursor & Antigravity.
origin/main
81/81 Tests Passing (100%)
v2.2.0
Aug 21, 2026
Enterprise Fleet • Unified Compute Provider Hub & 3-Card Onboarding: Seamless switching between Shadeform (20+ clouds), AWS Direct Spot ($0.75/hr), and Local Homelab ($0.00/hr).
• Automatic Secret Redaction: Masks API keys (********) on GET requests to prevent credential leaks.
• Fleet Availability Matrix: Live pricing comparison table in Cockpit with ⭐ g6e.xlarge (L40S 48GB) gold standard badge.
• Section 2.2 Universal AI Distribution Mesh: Integration specifications for LiteLLM, FastMCP, and open Agent Skills (SKILL.md).
e7098cf
63/63 Tests Passing
v2.1.0
Aug 21, 2026
Creator & Safety • 3D Camera Trajectory Compass: Interactive Pan, Tilt, Zoom, Roll dials with dynamic STG guidance scaling.
• Dead Man's Switch (Auto-Shutdown): Non-blocking background watchdog terminating idle GPU instances after configurable threshold (10m–60m).
• Dynamic Compute Quote: Real-time quoting switching between $0.00 Local Mock and ~$0.04 AWS Spot based on fleet state.
• Live 1s Cost Odometer: Absolute AWS LaunchTime odometer ticker in sidebar and cockpit.
ccbbb6a
61/61 Tests Passing
v2.0.0
Aug 20, 2026
GA Release • Resident LTX-2.5 Quantized Float8 Daemon: 0.0s cold start on 48GB VRAM.
• In-Process Kokoro-82M TTS Engine: Broadcast voice generation with LUFS normalization.
• Path Traversal & Security Boundary: resolve_output() regex sandboxing and 50MB binary limits.
• Zero-Build Frontend Architecture: Vanilla ES6 modules + native CSS variables.
55fa43b
55/55 Tests Passing