Metadata-Version: 2.5
Name: langchain-getyoutubetranscript
Version: 0.1.0
Summary: YouTube transcript tools and document loader for LangChain: get YouTube video transcripts with timestamps and search YouTube from agents and RAG pipelines, via the GetYouTubeTranscript API.
Project-URL: Homepage, https://getyoutubetranscript.com
Project-URL: Documentation, https://github.com/tubeagentkit/langchain-getyoutubetranscript#readme
Project-URL: Repository, https://github.com/tubeagentkit/langchain-getyoutubetranscript
Project-URL: API reference, https://getyoutubetranscript.com/docs
Project-URL: Get an API key, https://getyoutubetranscript.com/dashboard
Author: tubeagentkit
License: MIT
License-File: LICENSE
Keywords: agents,captions,document-loader,langchain,llm,rag,subtitles,tools,transcript,youtube,youtube-transcript
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Multimedia :: Video
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Requires-Dist: getyoutubetranscript<1.0.0,>=0.3.0
Requires-Dist: langchain-core<2.0.0,>=0.3.0
Provides-Extra: test
Requires-Dist: langchain-tests>=0.3.0; extra == 'test'
Requires-Dist: pytest>=7.0; extra == 'test'
Description-Content-Type: text/markdown

# langchain-getyoutubetranscript

YouTube transcript tools and a document loader for [LangChain](https://www.langchain.com), powered by the [GetYouTubeTranscript](https://getyoutubetranscript.com) API. Give agents the transcript of any YouTube video (optionally with `[m:ss]` timestamps) and YouTube search, or load transcripts into RAG pipelines.

The API fetches transcripts on its own servers, so it works from cloud servers and serverless functions without proxies, and without `RequestBlocked` / `IpBlocked` errors.

## Install

```bash
pip install langchain-getyoutubetranscript
```

Get an API key at [getyoutubetranscript.com/dashboard](https://getyoutubetranscript.com/dashboard) (free tier included) and set it:

```bash
export GETYOUTUBETRANSCRIPT_API_KEY=sk_live_...
```

Or pass `api_key="..."` to any class.

## Tools

| Class | Tool name | What it does |
| --- | --- | --- |
| `GetYouTubeTranscriptTool` | `youtube_transcript` | Transcript of a video (URL or ID) with title and channel. `timestamps=True` prefixes each line with `[m:ss]`. |
| `GetYouTubeTranscriptSearchTool` | `youtube_search` | Search YouTube for videos or channels. Returns JSON with a `continuation_token` for the next page. |

```python
from langchain_getyoutubetranscript import GetYouTubeTranscriptTool

tool = GetYouTubeTranscriptTool()
print(tool.invoke({"video": "https://youtu.be/jNQXAC9IVRw", "timestamps": True}))
# Title: Me at the zoo
# Channel: jawed
# Language: en
#
# [0:01] All right, so here we are, in front of the elephants
# ...
```

API errors (no captions, invalid video, out of credits) come back to the agent as a short message instead of raising.

### With an agent

```python
from langchain.agents import create_agent
from langchain_getyoutubetranscript import GetYouTubeTranscriptSearchTool, GetYouTubeTranscriptTool

agent = create_agent(
    model="anthropic:claude-sonnet-4-5",
    tools=[GetYouTubeTranscriptSearchTool(), GetYouTubeTranscriptTool()],
)
agent.invoke(
    {"messages": [{"role": "user", "content": "Find a short talk on transformers and summarize it with timestamps."}]}
)
```

## Document loader

```python
from langchain_getyoutubetranscript import GetYouTubeTranscriptLoader

docs = GetYouTubeTranscriptLoader(
    ["https://youtu.be/jNQXAC9IVRw", "5e37ZT3SQbk"],
    language="en",
    timestamps=False,
).load()

print(docs[0].metadata)
# {'source': 'https://www.youtube.com/watch?v=jNQXAC9IVRw', 'video_id': 'jNQXAC9IVRw',
#  'title': 'Me at the zoo', 'author_name': 'jawed', 'language_code': 'en', 'word_count': 39}
```

One `Document` per video. With `timestamps=True`, `page_content` is `[m:ss]` lines and `metadata["segments"]` holds `{start, duration, text}` per caption line (seconds).

## Pricing

Each transcript or search request uses one credit from your GetYouTubeTranscript account. Failed requests are not charged.

## Development

```bash
pip install -e ".[test]"
pytest                                        # unit + LangChain standard tests, no network
GETYOUTUBETRANSCRIPT_API_KEY=... pytest tests/integration_tests   # live, spends credits
```

## Links

- [GetYouTubeTranscript API docs](https://getyoutubetranscript.com/docs)
- [Python SDK](https://pypi.org/project/getyoutubetranscript/) (this package is built on it)
- [License: MIT](LICENSE)
