Metadata-Version: 2.5
Name: boson-video
Version: 0.2.0
Summary: Skim any video: a timed summary, a bilingual transcript, screen text and checked answers, as a web page or an MCP plugin.
Project-URL: Homepage, https://github.com/linboxin/Boson-Video
Project-URL: Repository, https://github.com/linboxin/Boson-Video
Project-URL: Issues, https://github.com/linboxin/Boson-Video/issues
Author: James Lin
License-Expression: MIT
License-File: LICENSE
Keywords: claude,learning,mcp,mcp-server,ocr,subtitles,summary,transcript,video,youtube
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: End Users/Desktop
Classifier: Natural Language :: Chinese (Simplified)
Classifier: Natural Language :: English
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Multimedia :: Video
Classifier: Topic :: Text Processing
Requires-Python: >=3.12
Requires-Dist: deno>=2.0
Requires-Dist: httpx[http2]>=0.28
Requires-Dist: imageio-ffmpeg>=0.6.0
Requires-Dist: mcp>=2.3.0
Requires-Dist: numpy>=2.0
Requires-Dist: pillow>=11.0
Requires-Dist: rapidocr-onnxruntime>=1.4.4
Requires-Dist: sherpa-onnx-core>=1.13.8
Requires-Dist: sherpa-onnx>=1.13.8
Requires-Dist: typesafe-sdk>=0.7.2
Requires-Dist: yt-dlp[default]>=2026.8.19
Description-Content-Type: text/markdown

# Boson-Video

**Get the point of any video without watching all of it.** Paste a YouTube link or drop a video
file. Every line of the summary, the transcript and the screen text carries the second it came
from, and questions are answered from the video and checked against it.

<!-- mcp-name: io.github.linboxin/boson-video -->

![A video read in Boson-Video: the player and the ribbon on the left, the terms explained on the right](https://raw.githubusercontent.com/linboxin/Boson-Video/main/docs/screenshot.jpg)

- **The ribbon:** the whole video on one strip (scenes, chapters, most replayed). Hover to see any second's frame and words.
- **Summary:** sections with their frames; every sentence timed and checked (✓ ? ✗).
- **Transcript:** the original with English beneath, technical terms explained, on-screen text read in.
- **Ask:** answers that cite their moments and show the frames, with background kept apart.
- **Your own AI:** the same document as an MCP plugin for Claude, Cursor, Codex and more.

Built for long talks, lectures and finance videos in a language you half know (Chinese and English today).

## In your own AI

One line, nothing else to install ([uv](https://docs.astral.sh/uv/) runs it):

```bash
claude mcp add boson-video -- uvx boson-video mcp
```

Claude Desktop (`claude_desktop_config.json`) or Cursor (`.cursor/mcp.json`):

```json
{ "mcpServers": { "boson-video": { "command": "uvx", "args": ["boson-video", "mcp"] } } }
```

Codex (`~/.codex/config.toml`):

```toml
[mcp_servers.boson-video]
command = "uvx"
args = ["boson-video", "mcp"]
```

Then ask your AI about any YouTube link. It opens the video, reads the transcript, looks at the
frames that matter at full resolution, searches, and checks its claims, citing every moment. No
keys needed. Everything runs on your computer. Details: [docs/plugin.md](https://github.com/linboxin/Boson-Video/blob/main/docs/plugin.md).

## Run the page

```bash
uvx boson-video web             # http://127.0.0.1:8770; the first run prints your invite code
```

Keys make it fuller: `INCEPTION_API_KEY` (Mercury writes the summary, the English and the terms)
and `TYPESAFE_API_KEY` (Jev checks every sentence and answers questions). Without them you still get
the scenes, the transcript and search.

## Tested videos

Measured, one run each unless a range is shown. Mac: M5 MacBook with Apple's transcriber. Windows: a 12-thread laptop with
SenseVoice or Parakeet on the CPU. Full tables: [docs/measured.md](https://github.com/linboxin/Boson-Video/blob/main/docs/measured.md).

| Video | Length | Scene map | Words | Transcript errors¹ | Screen text | Summary ✓ |
| --- | --- | --- | --- | --- | --- | --- |
| Money or Life 美股频道, Meta and AI (zh, talking head) | 24:53 | 1.3 s | 13.8 s (Mac) | no human captions | — | 13 / 13 |
| Money or Life 美股频道, AI drug discovery (zh, slides) | 28:24 | 1.1 s | 18.5 s (Mac) | not scored | 36 moments read | 15 / 16 |
| 程序员老王, LLM abliteration (zh, animated diagrams) | 11:35 | — | 10.0 s (Windows) | not scored | 88 moments; subtitles apart at 85 | 21–22 / 22–23 (several runs) |
| 陳永儀, TEDxTaipei (zh, talk) | 14:29 | — | 7.1 s (Mac) | 5.3% chars (Mac), 4.8% (SenseVoice) | — | 14 / 14 |
| Ken Robinson, TED (en, talk) | 20:06 | — | 9.9 s (Mac) | 10.0% words (Mac), 10.5% (Parakeet) | — | 22 / 25 |
| Sean's AI Stories, agent observability (en, screen recording) | 20:48 | 1.5 s | 13.2 s (Mac) | not scored | none: YouTube refused full resolution | 23 / 26 |
| Andrej Karpathy, Intro to LLMs (en, slides) | 59:48 | 1.6 s | 30.5 s (Mac) | only automatic captions | 19 of 20 slides right | 22 / 23 |
| Rick Astley, Never Gonna Give You Up (fast cuts) | 3:33 | 0.7 s | — | — | — | — |

¹ Against human captions (`scripts/accuracy.py`); about half the English errors are filler words the captions leave out.
Summary ✓: sentences confirmed against what was said (Jev and code).

## Command line

```bash
uvx boson-video <link | id | file>        # build a video: scenes, words, screen text, summary
uvx boson-video ask <video> "question"    # the moment that answers it
uvx boson-video models                    # fetch the local speech models now instead of on first use
```

From a clone: `uv sync`, then `uv run boson-video …`; `uv run pytest` runs the tests (offline).

## More

- [docs/workflow.md](https://github.com/linboxin/Boson-Video/blob/main/docs/workflow.md): how the workflow is designed: stages, the document, the ways in
- [docs/how-it-works.md](https://github.com/linboxin/Boson-Video/blob/main/docs/how-it-works.md): the pipeline, what YouTube allows, limits
- [docs/DIRECTION.md](https://github.com/linboxin/Boson-Video/blob/main/docs/DIRECTION.md): what we're building, milestones, decisions
- [docs/timeline-format.md](https://github.com/linboxin/Boson-Video/blob/main/docs/timeline-format.md): the document format, for other tools
- [AGENTS.md](https://github.com/linboxin/Boson-Video/blob/main/AGENTS.md): the guide for coding agents

MIT licensed. YouTube's terms don't allow automated access for products, so everything that touches
YouTube runs on your own computer.
