MCP YouTube Intelligence
Integrates with Google Gemini models to provide AI-powered video reporting and summarization.
Enables local, offline video analysis and summarization by connecting to Ollama instances running various LLMs.
Integrates with OpenAI models like GPT-4o-mini for intelligent video summarization and content processing.
Supports using PostgreSQL as a persistent database backend for caching and searching video data.
Allows for monitoring YouTube channels for new content updates through integrated RSS feed tracking.
Provides local storage and caching for video metadata, transcripts, and analysis results to optimize performance.
Enables searching for videos, fetching metadata, extracting transcripts, analyzing comments, and processing playlists to provide comprehensive video intelligence.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP YouTube IntelligenceSummarize this video and analyze the viewer sentiment: https://youtu.be/LV6Juz0xcrY"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
π English | νκ΅μ΄
MCP YouTube Intelligence
YouTube μμμ μ§λ₯μ μΌλ‘ λΆμνλ MCP μλ² + CLI
MCP (Model Context Protocol)λ Claude, Cursor κ°μ AI λκ΅¬κ° μΈλΆ μλΉμ€λ₯Ό μ¬μ©ν μ μκ² ν΄μ£Όλ νμ€ νλ‘ν μ½μ λλ€. μ΄ μλ²λ₯Ό μ°κ²°νλ©΄ "μ΄ μμ μμ½ν΄μ€" νλ§λλ‘ λΆμμ΄ μλ£λ©λλ€.
π― ν΅μ¬ κ°μΉ: μλ³Έ μλ§(2,000~30,000 ν ν°)μ μλ²μμ μ²λ¦¬νμ¬ LLMμλ ~200β500 ν ν°λ§ μ λ¬ν©λλ€.
π€ μ μ΄ μλ²μΈκ°?
λλΆλΆμ YouTube MCP μλ²λ μλ³Έ μλ§μ κ·Έλλ‘ LLMμ λμ§λλ€.
κΈ°λ₯ | κΈ°μ‘΄ MCP μλ² | MCP YouTube Intelligence |
μλ§ μΆμΆ | β | β |
μλ²μ¬μ΄λ μμ½ (ν ν° μ΅μ ν) | β | β |
ꡬ쑰νλ 리ν¬νΈ (μμ½+ν ν½+μν°ν°+λκΈ) | β | β |
μ±λ λͺ¨λν°λ§ (RSS) | β | β |
λκΈ κ°μ± λΆμ | β | β |
ν ν½ μΈκ·Έλ©ν μ΄μ | β | β |
μν°ν° μΆμΆ (ν/μ 200+κ°) | β | β |
μλ§/YouTube κ²μ | β | β |
λ°°μΉ μ²λ¦¬ | β | β |
SQLite/PostgreSQL μΊμ | β | β |
Related MCP server: yt-fetch
π λΉ λ₯Έ μμ
1. μ€μΉ
pip install mcp-youtube-intelligence
pip install yt-dlp # μλ§ μΆμΆμ νμπ‘ LLM μμ΄λ κΈ°λ³Έ μμ½(ν΅μ¬ λ¬Έμ₯ μΆμΆ)μ λμν©λλ€. κ³ νμ§ μμ½μ μνλ©΄ μλ LLM μ€μ μ μ°Έκ³ νμΈμ.
2. 첫 λ²μ§Έ λͺ λ Ήμ΄ μ€ν
# 리ν¬νΈ μμ± β μμ½, ν ν½, μν°ν°, λκΈμ νλ²μ λΆμ (LLM μ°λνμ)
mcp-yt report "https://www.youtube.com/watch?v=LV6Juz0xcrY"
# μλ§ μμ½λ§
mcp-yt transcript "https://www.youtube.com/watch?v=LV6Juz0xcrY"
# μμ IDλ§ μ¨λ λ©λλ€
mcp-yt report LV6Juz0xcrYβ οΈ zsh μ¬μ©μ: URLμ
?κ° μμΌλ―λ‘ λ°λμ λ°μ΄νλ‘ κ°μΈμΈμ.
π 리ν¬νΈ μΆλ ₯ μμ
mcp-yt report "https://www.youtube.com/watch?v=LV6Juz0xcrY" μ€ν κ²°κ³Ό (extractive μμ½):
# πΉ Video Analysis Report: OpenClaw Use Cases that are Actually Helpful! (ClawdBot)
> Channel: Duncan Rogoff | AI Automation | Duration: 16:29 | Language: en_ytdlp
## 1. Summary
OpenClaw is the most powerful AI agent framework in the world right now and
it's about to replace your entire workflow. I spent over $200 in the last
48 hours stress testing the system so you don't have to. It defines who it
is, how it behaves, and crucial behavioral boundaries. If you think open
claw is cool, just check out this video up here of 63 insane use cases
that other people are doing.
## 2. Key Topics
| # | Topic | Keywords | Timespan |
|---|-------|----------|----------|
| 1 | framework, world, right | framework, world, right | 0:00~0:05 |
| 2 | like, really, there | like, really, there | 0:05~2:23 |
| 3 | like, max, using | like, max, using | 2:23~4:22 |
| 4 | going, like, something | going, like, something | 4:22~5:03 |
| 5 | like, agents, basically | like, agents, basically | 5:03~6:04 |
| ... | ... | ... | ... |
| 15 | think, open, claw | think, open, claw | 16:24~16:29 |
## 4. Keywords & Entities
- **Technology**: GitHub, LLM, GPT
- **Company**: Anthropic, Apple
## 5. Viewer Reactions
- Total comments: 20
- Sentiment: Positive 45% / Negative 0% / Neutral 55%
- Top opinions:
- **@geetee2583** (positive, π8): Great info. Just need your inset video out of the way...
- **@bdog4026** (positive, π3): This tool is wild! Definitely the most in depth explanation...
- **@magalyvilela4917** (neutral, π3): Came to this video wondering it gonna teach me how to set up...π CLI μ 체 λͺ λ Ήμ΄
π 리ν¬νΈ (ν΅μ¬ κΈ°λ₯)
β οΈ **리ν¬νΈμ μμ½ μΉμ μ LLM μ°λμ΄ νμμ λλ€. Ollama λΉ λ₯Έ μ€μ (무λ£, 3λΆμ΄λ©΄ λ):
# 1. Ollama μ€μΉ: https://ollama.ai # 2. λͺ¨λΈ λ€μ΄λ‘λ ollama pull qwen2.5:7b # 3. νκ²½λ³μ μ€μ export MYI_LLM_PROVIDER=ollama export MYI_OLLAMA_MODEL=qwen2.5:7b # μ격 μλ²λΌλ©΄ νΈμ€νΈλ μ§μ export MYI_OLLAMA_BASE_URL=http://your-server:11434
mcp-yt report "https://youtube.com/watch?v=VIDEO_ID"
mcp-yt report VIDEO_ID --provider ollama # LLM νλ‘λ°μ΄λ μ§μ
mcp-yt report VIDEO_ID --no-comments # λκΈ μ μΈ
mcp-yt report VIDEO_ID -o report.md # νμΌ μ μ₯π― μλ§ μΆμΆ + μμ½
mcp-yt transcript VIDEO_ID # μμ½ (~200β500 ν ν°)
mcp-yt transcript VIDEO_ID --mode full # μ 체 μλ§
mcp-yt transcript VIDEO_ID --mode chunks # μ²ν¬ λΆν
mcp-yt --json transcript VIDEO_ID # JSON μΆλ ₯κΈ°ν
mcp-yt video VIDEO_ID # λ©νλ°μ΄ν°
mcp-yt comments VIDEO_ID --max 20 # λκΈ (κ°μ± λΆμ ν¬ν¨)
mcp-yt entities VIDEO_ID # μν°ν° μΆμΆ
mcp-yt segments VIDEO_ID # ν ν½ μΈκ·Έλ©ν
μ΄μ
mcp-yt search "ν€μλ" --max 5 # YouTube κ²μ
mcp-yt monitor subscribe @μ±λνΈλ€ # μ±λ λͺ¨λν°λ§
mcp-yt playlist PLAYLIST_ID # νλ μ΄λ¦¬μ€νΈ
mcp-yt batch ID1 ID2 ID3 # λ°°μΉ μ²λ¦¬
mcp-yt search-transcripts "ν€μλ" # μ μ₯λ μλ§ κ²μπ‘ λͺ¨λ λͺ λ Ήμ΄μ
--jsonνλκ·Έλ₯Ό μΆκ°νλ©΄ JSON μΆλ ₯λ©λλ€.
π MCP μλ² μ°κ²°
MCP μλ²λ stdio νλ‘ν μ½λ‘ ν΅μ ν©λλ€.
Claude Desktop / Cursor / OpenCode
μ€μ νμΌμ μΆκ° (claude_desktop_config.json, .cursor/mcp.json, mcp.json):
{
"mcpServers": {
"youtube": {
"command": "uvx",
"args": ["mcp-youtube-intelligence"],
"env": {
"MYI_LLM_PROVIDER": "ollama",
"MYI_OLLAMA_MODEL": "qwen2.5:7b"
}
}
}
}π‘
uvxλuvν¨ν€μ§ λ§€λμ μ μ€ν λͺ λ Ήμ΄μ λλ€.pip install uvλ‘ μ€μΉνμΈμ.ν΄λΌμ°λ LLMμ μ°λ €λ©΄
envμ API ν€λ₯Ό μΆκ°νλ©΄ λ©λλ€:"OPENAI_API_KEY": "sk-..."
Claude Code
claude mcp add youtube -- uvx mcp-youtube-intelligenceMCP Tools (9κ°)
Tool | μ€λͺ | μμ ν ν° |
| λ©νλ°μ΄ν° + μμ½ | ~200β500 |
| μλ§ (summary/full/chunks) | ~200β500 |
| λκΈ + κ°μ± λΆμ | ~200β500 |
| RSS μ±λ λͺ¨λν°λ§ | ~100β300 |
| μ μ₯λ μλ§ κ²μ | ~100β400 |
| μν°ν° μΆμΆ | ~150β300 |
| ν ν½ λΆν | ~100β250 |
| YouTube κ²μ | ~200 |
| νλ μ΄λ¦¬μ€νΈ λΆμ | ~200β500 |
get_video
νλΌλ―Έν° | νμ | νμ | μ€λͺ |
| string | β | YouTube μμ ID |
get_transcript
νλΌλ―Έν° | νμ | νμ | κΈ°λ³Έκ° | μ€λͺ |
| string | β | β | YouTube μμ ID |
| string | β |
|
|
get_comments
νλΌλ―Έν° | νμ | νμ | κΈ°λ³Έκ° | μ€λͺ |
| string | β | β | YouTube μμ ID |
| int | β |
| λ°νν λκΈ μ |
| bool | β |
| μμ½ λ·° |
monitor_channel
νλΌλ―Έν° | νμ | νμ | κΈ°λ³Έκ° | μ€λͺ |
| string | β | β | μ±λ URL/@νΈλ€/ID |
| string | β |
|
|
search_transcripts
νλΌλ―Έν° | νμ | νμ | κΈ°λ³Έκ° | μ€λͺ |
| string | β | β | κ²μ ν€μλ |
| int | β |
| μ΅λ κ²°κ³Ό μ |
extract_entities / segment_topics
νλΌλ―Έν° | νμ | νμ | μ€λͺ |
| string | β | YouTube μμ ID |
search_youtube
νλΌλ―Έν° | νμ | νμ | κΈ°λ³Έκ° | μ€λͺ |
| string | β | β | κ²μ ν€μλ |
| int | β |
| μ΅λ κ²°κ³Ό μ |
| string | β |
|
|
get_playlist
νλΌλ―Έν° | νμ | νμ | κΈ°λ³Έκ° | μ€λͺ |
| string | β | β | νλ μ΄λ¦¬μ€νΈ ID |
| int | β |
| μ΅λ μμ μ |
βοΈ μ€μ
LLM νλ‘λ°μ΄λ μ€μ
LLM μμ΄λ κΈ°λ³Έ μμ½(ν΅μ¬ λ¬Έμ₯ μΆμΆ)μ λμν©λλ€. κ³ νμ§ μμ½μ μνλ©΄:
Ollama (μΆμ² β 무λ£, μ€νλΌμΈ)
# 1. Ollama μ€μΉ: https://ollama.ai
# 2. λͺ¨λΈ λ€μ΄λ‘λ
ollama pull qwen2.5:7b
# 3. νκ²½λ³μ μ€μ
export MYI_LLM_PROVIDER=ollama
export MYI_OLLAMA_MODEL=qwen2.5:7b
# 4. (μ ν) μ격 Ollama μλ² μ¬μ© μ
export MYI_OLLAMA_BASE_URL=http://your-server:11434ν΄λΌμ°λ LLM
# API ν€λ§ μ€μ νλ©΄ μλ κ°μ§ (MYI_LLM_PROVIDER=auto)
export OPENAI_API_KEY=sk-... # OpenAI
export ANTHROPIC_API_KEY=sk-ant-... # Anthropic
export GOOGLE_API_KEY=AIza... # Google
# νΉμ νλ‘λ°μ΄λ μ§μ
export MYI_LLM_PROVIDER=anthropicν΄λΌμ°λ LLM ν¨ν€μ§:
pip install "mcp-youtube-intelligence[llm]"(OpenAI) /[anthropic-llm]/[google-llm]/[all-llm]
μΆμ² Ollama λͺ¨λΈ
λͺ©μ | λͺ¨λΈ | ν¬κΈ° | νκ΅μ΄ | μμ΄ | νμ§ |
λ€κ΅μ΄ (μΆμ²) |
| 4.4GB | β | β | βββ |
μμ΄ μ€μ¬ |
| 4.7GB | β οΈ | β | βββ |
νκ΅μ΄ νΉν |
| 5.4GB | β | β | βββ |
κ²½λ |
| 1.9GB | β | β | ββ |
λ€κ΅μ΄ νΉν |
| 4.8GB | β | β | βββ |
β±οΈ μ€μΈ‘ λ²€μΉλ§ν¬
RTX 3070 8GB Β· Ollama Β· νκ΅μ΄ μλ§ ~2,900μ (5λΆ 19μ΄ μμ)
load_durationμ μΈ, μμ μμ± μκ° κΈ°μ€
λͺ¨λΈ | Prompt μ²λ¦¬ | μμ± μκ° | μλ | μΆλ ₯ | νμ§ |
Extractive | - | μ¦μ | - | 379μ | ββ |
qwen2.5:1.5b | 7.8s | 4.7s | 30.4 tok/s | 232μ | ββ |
qwen2.5:7b | 34.5s | 18.8s | 7.3 tok/s | 766μ | βββ |
aya-expanse:8b | 29.5s | 34.5s | 6.2 tok/s | 405μ | βββ |
β οΈ μ²« μ€ν μ λͺ¨λΈ λ‘λμ 15~60μ΄ μΆκ°.
keep_aliveλ‘ λ©λͺ¨λ¦¬ μ μ§νλ©΄ μ΄ν λ‘λ μμ.
νκ²½λ³μ | κΈ°λ³Έκ° | μ€λͺ |
|
| λ°μ΄ν° λλ ν 리 |
|
|
|
|
| SQLite κ²½λ‘ |
| β | PostgreSQL DSN |
|
| yt-dlp κ²½λ‘ |
|
| μ΅λ λκΈ μ |
|
|
|
| β | OpenAI ν€ |
|
| OpenAI λͺ¨λΈ |
| β | Anthropic ν€ |
|
| Anthropic λͺ¨λΈ |
| β | Google ν€ |
|
| Google λͺ¨λΈ |
|
| Ollama URL |
|
| Ollama λͺ¨λΈ |
|
| vLLM URL |
| β | vLLM λͺ¨λΈ |
|
| LM Studio URL |
| β | LM Studio λͺ¨λΈ |
π νΈλ¬λΈμν
λ¬Έμ | ν΄κ²° |
| URLμ λ°μ΄νλ‘ κ°μΈκΈ°: |
|
|
μλ§ μλ μμ |
|
SQLite database locked | μλ² μΈμ€ν΄μ€ νλλ§ μ€ν μ€μΈμ§ νμΈ |
LLM μμ½ μ€ν¨ | μλμΌλ‘ extractive ν΄λ°±λ¨. API ν€ νμΈ. |
π€ Contributing
git clone https://github.com/JangHyuckYun/mcp-youtube-intelligence.git
cd mcp-youtube-intelligence
pip install -e ".[dev]"
pytest tests/ -vπ λΌμ΄μ μ€
Apache 2.0 β LICENSE
π λ³κ²½ μ΄λ ₯
λ μ§ | λ²μ | μ£Όμ λ³κ²½ |
2025-02-18 | v0.1.0 | μ΄κΈ° λ¦΄λ¦¬μ€ β 9κ° MCP λꡬ, CLI, SQLite |
2025-02-18 | v0.1.1 | Multi-LLM (OpenAI/Anthropic/Google), Apache 2.0 |
2025-02-18 | v0.1.2 | Local LLM (Ollama/vLLM/LM Studio), yt-dlp μλ§ κ°μ , μμ΄ κΈ°λ³Έ μΆλ ₯ |
2025-02-18 | v0.1.3 | Local LLM (Ollama/vLLM/LM Studio), yt-dlp μλ§ κ°μ , μμ΄ κΈ°λ³Έ μΆλ ₯ |
Available Tools
10 toolsextract_entitiesC
Extract structured entities (companies, indices, people, sectors, etc.) from a video transcript.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | YouTube video ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool extracts entities but doesn't describe how (e.g., via NLP, accuracy, rate limits), what the output looks like (since no output schema exists), or any constraints (e.g., video length limits, processing time). This leaves significant gaps for an AI agent to understand the tool's behavior beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that clearly states the tool's purpose without unnecessary words. It is front-loaded with the core action and resource, making it easy to parse. Every part of the sentence contributes essential information, earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of entity extraction (which involves NLP processing) and the lack of annotations and output schema, the description is incomplete. It doesn't explain the output format, accuracy, limitations, or how it integrates with other tools (e.g., needing 'get_transcript' first). For a tool with no structured behavioral data, more context is needed to guide effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the single parameter 'video_id' documented as a YouTube video ID. The description adds no additional parameter semantics beyond what the schema provides, such as format examples or validation rules. However, with high schema coverage and only one parameter, the baseline score of 3 is appropriate as the schema adequately covers parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: extracting structured entities from video transcripts. It specifies the resource (video transcripts) and the action (extracting entities), and lists the types of entities (companies, indices, people, sectors, etc.). However, it does not explicitly differentiate this tool from sibling tools like 'segment_topics' or 'search_transcripts', which might also process transcripts, so it doesn't fully distinguish from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a transcript first), exclusions, or compare it to siblings like 'segment_topics' or 'search_transcripts' that might handle similar data. Usage is implied by the purpose but lacks explicit context or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_reportB
Generate a structured markdown report for a YouTube video. Includes summary, topic segments, entities, and optionally comments.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | YouTube video ID | |
| include_comments | No | Include comment analysis | |
| llm_provider | No | LLM provider for summary |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions the report includes summary, topic segments, entities, and optionally comments, but lacks details on behavioral traits like processing time, error handling, authentication needs, or rate limits. For a tool with no annotations, this is a significant gap in transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the main purpose and lists key components. It avoids unnecessary details, though it could be slightly more structured for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 3 parameters with full schema coverage, the description is adequate but incomplete. It covers the purpose and components but lacks behavioral context and output details, making it minimally viable for a report-generation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters (video_id, include_comments, llm_provider). The description adds minimal value by implying the report uses an LLM for summary and includes comment analysis, but doesn't provide additional syntax or format details beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'generate' and the resource 'structured markdown report for a YouTube video', specifying the content includes summary, topic segments, entities, and optionally comments. It distinguishes from siblings like get_transcript or extract_entities by focusing on report generation, though it doesn't explicitly contrast with all alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when a comprehensive report is needed, mentioning optional comment inclusion, but doesn't specify when to use this versus siblings like get_video or segment_topics individually. No explicit alternatives or exclusions are provided, leaving some ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_commentsC
Get top comments for a video. Optionally summarize them.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | YouTube video ID | |
| top_n | No | Number of top comments | |
| summarize | No | Return summarized view |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool gets top comments and optionally summarizes them, but it doesn't reveal critical behavioral traits such as whether this is a read-only operation, potential rate limits, authentication needs, error handling, or what 'top' means (e.g., by likes, recency). This leaves significant gaps for an AI agent to understand how the tool behaves in practice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded, consisting of just one sentence that directly states the tool's purpose and optional feature. Every word earns its place without any redundancy or unnecessary elaboration, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (3 parameters, no output schema, no annotations), the description is incomplete. It lacks details on behavioral aspects like safety, performance, or output format, and doesn't compensate for the absence of annotations or output schema. For a tool that retrieves and potentially summarizes comments, more context is needed to ensure the agent can use it effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, clearly documenting all parameters (video_id, top_n, summarize) with their types and defaults. The description adds minimal value beyond this, only implying the optional summarization feature, which is already covered in the schema. According to the rules, with high schema coverage (>80%), the baseline is 3 even with no additional param info in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get top comments for a video' specifies the verb ('Get') and resource ('top comments for a video'), making it understandable. However, it doesn't explicitly differentiate from sibling tools like 'get_video' or 'get_transcript', which might also involve video-related data retrieval, so it lacks sibling differentiation for a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It mentions an optional summarization feature but doesn't explain when to use 'summarize' or how this tool compares to siblings like 'search_transcripts' or 'get_video' for video analysis tasks. Without any usage context or exclusions, it falls short of providing helpful guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_playlistC
Get playlist metadata and video list from a YouTube playlist.
| Name | Required | Description | Default |
|---|---|---|---|
| playlist_id | Yes | YouTube playlist ID (e.g. PLrAXtmErZgOeiKm4sgNOknGvNjby9efdf) | |
| max_videos | No | Max videos to retrieve |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool retrieves metadata and video lists, implying a read-only operation, but doesn't cover important aspects like rate limits, authentication needs, error handling, or pagination behavior. This leaves significant gaps in understanding how the tool behaves beyond its basic function.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without any unnecessary words or fluff. It is front-loaded with the core action and resources, making it highly concise and well-structured for quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (retrieving playlist data with two parameters), lack of annotations, and no output schema, the description is incomplete. It doesn't explain what metadata is included, the format of the video list, potential limitations, or error cases, leaving the agent with insufficient context for effective use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with clear documentation for both parameters ('playlist_id' and 'max_videos'), including examples and defaults. The description adds no additional parameter semantics beyond what the schema provides, so it meets the baseline of 3 for adequate but not enhanced parameter information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Get' and the resources 'playlist metadata and video list from a YouTube playlist', making the purpose specific and understandable. However, it doesn't explicitly distinguish this tool from sibling tools like 'get_video' or 'monitor_channel', which might also involve YouTube content retrieval, so it doesn't reach the highest score of 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any prerequisites, context for usage, or comparisons to sibling tools such as 'get_video' for individual videos or 'search_youtube' for broader searches, leaving the agent without explicit usage instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptA
Get video transcript. mode: 'summary' (default, ~300 tokens), 'full' (saves to file, returns path), 'chunks' (split into segments).
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | YouTube video ID | |
| mode | No | summary | |
| llm_provider | No | LLM provider for summary (default: auto) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses key behavioral traits: default mode, token length for summary (~300), file-saving behavior for 'full' mode, and segmentation for 'chunks'. However, it doesn't mention rate limits, authentication needs, error conditions, or what happens with invalid video IDs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise and front-loaded: the first three words state the core purpose, followed by efficient mode explanations. Every sentence earns its place by providing essential operational details without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no annotations and no output schema, the description provides adequate but incomplete context. It covers mode behaviors well but lacks information about return values (beyond 'returns path' for full mode), error handling, or performance characteristics that would help an agent use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67% (2 of 3 parameters have descriptions). The description adds significant value beyond the schema: it explains what each 'mode' does (summary length, file saving for full, segmentation for chunks) and clarifies the default behavior. This compensates well for the schema's partial coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get video transcript' with specific modes. It distinguishes from siblings like 'get_video' (metadata) and 'search_transcripts' (searching). However, it doesn't explicitly contrast with 'segment_topics' which might overlap with 'chunks' mode.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use different modes ('summary' for brief, 'full' for complete, 'chunks' for segmented), but doesn't provide explicit guidance on when to choose this tool over alternatives like 'search_transcripts' or 'get_video'. No exclusions or prerequisites are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_videoB
Get video metadata + summary (~300 tokens). Provide a YouTube video ID.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | YouTube video ID (e.g. dQw4w9WgXcQ) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the output includes 'metadata + summary (~300 tokens)', which gives some behavioral context about the response format and length. However, it doesn't address important aspects like whether this is a read-only operation, potential rate limits, authentication requirements, or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with just two sentences that each serve a clear purpose: first states what the tool does, second specifies the required input. There's zero wasted language or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read operation with no output schema, the description provides basic functionality but lacks important context. It doesn't explain what specific metadata fields are returned, how the summary is generated, or any limitations. The ~300 token mention is helpful but insufficient for full understanding of the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the single parameter completely. The description adds minimal value beyond the schema by specifying 'YouTube video ID' (implied in schema's example) and reinforcing it's required. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get video metadata + summary (~300 tokens)' with the specific resource being a YouTube video. It distinguishes from siblings like get_transcript (which gets transcript text) and get_comments (which gets comments), but doesn't explicitly mention these distinctions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides minimal usage guidance: 'Provide a YouTube video ID' indicates the required input but offers no context about when to use this tool versus alternatives like get_transcript or search_youtube. There's no mention of prerequisites, limitations, or comparative use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
monitor_channelB
Monitor a YouTube channel via RSS. action: 'add' (subscribe), 'check' (poll for new videos), 'list' (show subscriptions), 'remove' (unsubscribe).
| Name | Required | Description | Default |
|---|---|---|---|
| channel_ref | Yes | Channel URL, @handle, or ID | |
| action | No | check |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden but is insufficient. It mentions actions but doesn't disclose behavioral traits such as whether 'add' requires authentication, if 'check' polls at a specific rate, what 'list' returns, or if 'remove' is destructive. This leaves critical operational details unclear for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise and front-loaded: the first sentence states the core purpose, followed by a compact breakdown of actions. Every sentence earns its place with no wasted words, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (2 parameters, multiple actions), no annotations, and no output schema, the description is incomplete. It fails to explain return values, error conditions, or behavioral nuances like subscription persistence or polling intervals, which are essential for proper agent invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50% (channel_ref has a description, action does not). The description adds value by explaining action enum values (e.g., 'add' means subscribe), which compensates partially for the missing schema description for action. However, it doesn't clarify channel_ref formats beyond what the schema states, leaving gaps in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Monitor a YouTube channel via RSS' with specific actions. It distinguishes itself from siblings like get_video or search_youtube by focusing on RSS-based monitoring rather than direct API queries. However, it doesn't explicitly contrast with all siblings (e.g., get_playlist might also involve channel content).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through the action parameter breakdown (add, check, list, remove), suggesting when to use each mode. However, it lacks explicit guidance on when to choose this tool over alternatives like get_video for video retrieval or search_youtube for direct searches, and doesn't mention prerequisites like needing an RSS feed setup.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_transcriptsB
Search stored transcripts by keyword. Returns matching snippets.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search keyword or phrase | |
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It mentions the tool 'returns matching snippets,' which gives some output context, but lacks details on permissions, rate limits, error handling, or whether it's read-only/destructive. For a search tool with zero annotation coverage, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with zero wasteβit states the action and the result directly. It's appropriately sized and front-loaded, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (search with two parameters), no annotations, and no output schema, the description is minimally adequate. It covers the basic purpose and return type but lacks details on usage context, behavioral traits, and parameter nuances, leaving gaps for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50% (only 'query' has a description, 'limit' has none). The description adds no parameter semantics beyond what's in the schemaβit doesn't explain 'query' further or clarify 'limit' behavior (e.g., max value, pagination). With partial schema coverage, the description doesn't compensate adequately, meeting the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('search') and resource ('stored transcripts'), and mentions the return type ('matching snippets'). However, it doesn't explicitly differentiate from sibling tools like 'get_transcript' or 'search_youtube', which would require a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'get_transcript' (which might retrieve full transcripts) or 'search_youtube' (which might search YouTube content). It only states what the tool does, not when it's appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_youtubeB
Search YouTube videos by keyword. Returns metadata list (~200 tokens).
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search keyword or phrase | |
| max_results | No | Max results (1-50) | |
| channel_id | No | Limit search to a specific channel ID | |
| published_after | No | Filter: published after (ISO 8601) | |
| order | No | relevance |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions that the tool 'Returns metadata list (~200 tokens),' which gives some insight into output format and size, but lacks critical details like whether this is a read-only operation, rate limits, authentication requirements, or error handling. For a search tool with no annotations, this is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded: two sentences that directly state the tool's function and output. Every word earns its place, with no redundant information or fluff. It efficiently communicates the core purpose without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (5 parameters, no output schema, no annotations), the description is minimally adequate. It covers the basic purpose and output format but lacks details on behavioral traits, usage context, and parameter nuances. Without annotations or an output schema, the agent might struggle with full operational understanding, though the core function is clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 80% (high), so the baseline score is 3. The description adds minimal value beyond the schemaβit mentions 'by keyword,' which aligns with the 'query' parameter, but doesn't explain parameter interactions or provide additional context like search scope or result formatting. This meets the baseline but doesn't enhance understanding significantly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Search YouTube videos by keyword.' It specifies the verb ('Search') and resource ('YouTube videos'), making the function unambiguous. However, it doesn't explicitly differentiate this tool from sibling tools like 'search_transcripts' or 'get_video', which could cause confusion about when to use each.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention sibling tools like 'search_transcripts' (for searching within video transcripts) or 'get_video' (for retrieving specific video details), leaving the agent to infer usage context. There are no explicit when-to-use or when-not-to-use instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
segment_topicsC
Segment a video transcript into topics based on transition markers.
| Name | Required | Description | Default |
|---|---|---|---|
| video_id | Yes | YouTube video ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool segments transcripts but doesn't describe what 'transition markers' are, how topics are defined, the output format (e.g., list of segments with timestamps), error handling, or any rate limits. This leaves significant gaps for a tool that performs analysis on video content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that directly states the tool's purpose without unnecessary words. It's front-loaded with the core action and resource, making it easy to parse. Every part of the sentence earns its place by conveying essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of segmenting video transcripts (an analysis task) with no annotations and no output schema, the description is incomplete. It lacks details on behavioral traits, output format, and usage context, which are critical for an AI agent to invoke it correctly. The high schema coverage doesn't compensate for these gaps in a non-trivial tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, with the single parameter 'video_id' documented as 'YouTube video ID.' The description adds no additional meaning beyond this, such as format examples or constraints. With high schema coverage, the baseline score of 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('segment') and resource ('video transcript') with the specific purpose of dividing it 'into topics based on transition markers.' It distinguishes from siblings like 'get_transcript' (retrieval) or 'search_transcripts' (searching), but doesn't explicitly contrast with all alternatives. The purpose is specific but not fully differentiated from all siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites (e.g., needing a transcript first), exclusions, or compare to siblings like 'extract_entities' or 'search_transcripts' for similar text analysis tasks. Usage is implied by the purpose but lacks explicit context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
10 tool updates
v0.1.3- First observed
extract_entities - First observed
generate_report - First observed
get_comments - First observed
get_playlist - First observed
get_transcript - First observed
get_video - First observed
monitor_channel - First observed
search_transcripts - First observed
search_youtube - First observed
segment_topics
TDQS
Each tool has a clearly distinct purpose targeting specific YouTube-related tasks, such as extracting entities, generating reports, fetching comments, retrieving transcripts, and monitoring channels. There is no overlap in functionality, making it easy for an agent to select the correct tool without confusion.
All tool names follow a consistent verb_noun pattern using snake_case, such as 'extract_entities', 'generate_report', and 'get_transcript'. This uniformity enhances readability and predictability across the entire tool set.
With 10 tools, the server is well-scoped for YouTube intelligence tasks, covering key areas like video metadata, transcripts, comments, playlists, search, and monitoring. Each tool serves a unique and necessary function without being excessive or insufficient.
The tool set provides comprehensive coverage for YouTube video analysis, including data retrieval (video, transcript, comments), processing (entities, topics, reports), search (transcripts, YouTube), and monitoring (channel RSS). There are no apparent gaps that would hinder an agent's workflow.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
An MCP server that gives any LLM or agent clean YouTube transcripts on demand: a single video, a whole channel, or a playlist, plus AI cleanup of auto-generated captions. API-key auth, credit-based, same backend as the public v1 API. Get a free API key with 25 free credits at youtubetranscriptdownload.com/account.
MCP server for Google Veo AI video generation
Multimodal video analysis MCP β transcription, vision, and OCR for any video URL.
π― The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free.
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server that provides AI assistants with powerful tools to interact with YouTube, including video searching, transcript extraction, comment retrieval, and more.819Apache 2.0
- FlicenseNot gradedqualityDmaintenanceAn MCP server that enables interaction with the YouTube Data API, allowing users to search videos, get video and channel details, analyze trends, and fetch video transcripts.-
- AlicenseAqualityDmaintenanceAn MCP server that enables users to retrieve YouTube transcripts and perform video or channel searches without requiring Google API keys. It supports transcript chunking and provides tools for detailed video content analysis and channel metadata extraction.5584MIT
- FlicenseBqualityDmaintenanceAn MCP server that extracts transcripts, metadata, and summaries from YouTube videos across various URL formats including Shorts and standard links. It provides comprehensive video data and insights for analysis within MCP-compatible environments.3-
Appeared in Searches
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/JangHyuckYun/mcp-youtube-intelligence'
If you have feedback or need assistance with the MCP directory API, please join our Discord server