tubemcp
Search YouTube and fetch video transcripts.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@tubemcpsearch YouTube for machine learning tutorials"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
TubeMCP
MCP server that lets AI agents search YouTube and fetch transcripts. Zero config — just install and go.
What is MCP? Model Context Protocol lets AI assistants like Claude call external tools. TubeMCP gives your AI agent the ability to search YouTube and read any video's transcript — useful for summarization, Q&A, research, and content analysis.
Prerequisites
Python 3.10+ — download
Related MCP server: YouTube MCP Server
Installation
pip install tubemcpor
uv tool install tubemcpThen add it to your client:
Claude Code:
claude mcp add tubemcp -- tubemcpClaude Desktop — add to your claude_desktop_config.json:
{
"mcpServers": {
"tubemcp": {
"command": "tubemcp"
}
}
}Cursor — add to .cursor/mcp.json:
{
"mcpServers": {
"tubemcp": {
"command": "tubemcp"
}
}
}Windsurf — add to ~/.codeium/windsurf/mcp_config.json:
{
"mcpServers": {
"tubemcp": {
"command": "tubemcp"
}
}
}What you get
youtube_get_transcript
Fetch the English transcript and metadata for any YouTube video.
Input: A YouTube URL or video ID in any of these formats:
https://www.youtube.com/watch?v=VIDEO_IDhttps://youtu.be/VIDEO_IDhttps://www.youtube.com/embed/VIDEO_IDhttps://www.youtube.com/v/VIDEO_IDVIDEO_ID(bare 11-character ID)
Returns:
video_id— the video IDtitle— video titlechannel_name— channel namethumbnail_url— thumbnail URLduration_seconds— video durationpublish_date— publish datetranscript— full transcript textfrom_cache— whether the result was served from cache
youtube_search
Search YouTube with multiple queries for broader coverage. Results are deduplicated by video ID. Returns metadata only — no transcripts.
Input:
queries(list[str]) — search queries to run. Use 2–3 from different angles for best results.max_results_per_query(int, default3) — max results returned per query.
Returns a list of results, each containing:
video_id— the video IDtitle— video titlechannel_name— channel nameurl— video URLduration_seconds— video duration
Caching
Transcripts are cached locally in ~/.tubemcp/cache.db (SQLite). Subsequent requests for the same video are served instantly from cache.
Troubleshooting
spawn uvx ENOENT
This means your MCP client can't find the uvx command. Three fixes:
uv not installed — Install it: https://docs.astral.sh/uv/getting-started/installation/
uv not on PATH — Use the full path to
uvxin your config (find yours withwhich uvx):"command": "/Users/you/.local/bin/uvx"Switch to pip — Skip uv entirely. Install with
pip install tubemcpand use"command": "tubemcp"in your config (see pip installation above).
Verify uv is working:
uvx --versionDevelopment
git clone https://github.com/BlockBenny/tubemcp.git
cd tubemcp
pip install -e ".[dev]"
pytestContributing
See CONTRIBUTING.md for development setup and guidelines.
License
MIT
Available Tools
2 toolsyoutube_get_transcriptA
Fetch the transcript and metadata for a single YouTube video. Accepts a URL or video ID. Results are cached locally. Fetch only the videos most relevant to the user's question — avoid bulk fetching. The transcript field is an array of segments, each with text, start (seconds), and duration (seconds).
| Name | Required | Description | Default |
|---|---|---|---|
| video_url | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions local caching and the segment structure of transcripts, adding useful context. However, it does not state whether the operation is read-only, potential errors, or authentication requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each providing unique and necessary information: purpose, usage constraint, and result format. It is front-loaded and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with only one parameter and an output schema, the description covers the core purpose, usage context, parameter format, and even adds structural details about the transcript. It could mention error handling or rate limits, but these are not critical for a straightforward fetch operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only defines video_url as a string, but the description clarifies that it accepts a URL or video ID, which is essential for correct invocation. With 0% schema description coverage, this added meaning significantly helps the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool fetches transcript and metadata for a single YouTube video, using a URL or video ID. It distinguishes itself from the sibling tool youtube_search by specifying the exact resource (transcript) and action (fetch).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit guidance to 'fetch only the videos most relevant to the user's question — avoid bulk fetching,' which helps the agent decide when to use this tool. It does not explicitly contrast with youtube_search, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
youtube_searchA
Search YouTube with multiple queries for broader coverage. Each query runs a separate search; results are deduplicated by video ID. Use 2-3 queries from different angles for best results (e.g. a specific query, a broader one, and an alternative phrasing). Returns metadata only — no transcripts.
| Name | Required | Description | Default |
|---|---|---|---|
| queries | Yes | ||
| max_results_per_query | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of disclosing behavior. It reveals that each query runs a separate search, results are deduplicated by video ID, and only metadata is returned. While it doesn't mention pagination or rate limits, it covers the key behavioral traits that affect how an agent would use the results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-organized: it states the main capability, explains how queries are executed, offers usage guidance, and specifies the output scope. Every sentence contributes meaningful information with no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a search tool with only two parameters and an output schema present, the description is complete. It covers the purpose, usage strategy, behavioral nuances (deduplication, separate searches), and the nature of results (metadata only). There is no ambiguity about when or how to use this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description compensates by explaining the query strategy behind the 'queries' parameter, including typical count and phrasing. However, 'max_results_per_query' is not mentioned in the description, though its name and default value are self-explanatory. The guidance on query angles adds value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with 'Search YouTube with multiple queries for broader coverage,' clearly specifying the verb (Search), resource (YouTube), and distinguishing feature (multiple queries). This differentiates it from the sibling tool youtube_get_transcript, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly advises using 2-3 queries from different angles, providing concrete strategies (specific, broader, alternative phrasing). It also states 'Returns metadata only — no transcripts,' indicating when this tool is not appropriate, especially compared to the transcript-focused sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
2 tool updates
v0.1.3- First observed
youtube_get_transcript - First observed
youtube_search
TDQS
The two tools have completely distinct purposes: one searches for videos and returns metadata, the other fetches a transcript for a specific video. There is no overlap or ambiguity.
Both tools share the 'youtube_' prefix and use verb-like names, but 'youtube_get_transcript' follows a verb_object pattern while 'youtube_search' is just a verb. Minor deviation from a consistent pattern.
With only two tools, the server is minimal but covers its core workflow of search and transcript retrieval. This falls into the borderline range for tool count.
The search and get_transcript tools provide a complete workflow: an agent can find a video and then fetch its transcript. There are no obvious missing operations for the stated purpose.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
An MCP server that gives any LLM or agent clean YouTube transcripts on demand: a single video, a whole channel, or a playlist, plus AI cleanup of auto-generated captions. API-key auth, credit-based, same backend as the public v1 API. Get a free API key with 25 free credits at youtubetranscriptdownload.com/account.
MCP server for RiverScript, an AI transcription platform - fetches transcripts shared via a link.
An MCP server that gives your AI access to the source code and docs of all public github repos
Driflyte MCP server which lets AI assistants query topic-specific knowledge from web and GitHub.
Related MCP Servers
- AlicenseNot gradedqualityNot gradedmaintenanceAn MCP server that provides YouTube data access without API keys or quotas. It enables agents to search videos, retrieve transcripts and metadata, and perform full-text search across cached content for AI context retrieval.3-
- FlicenseNot gradedqualityDmaintenanceThis MCP server fetches and extracts transcripts from YouTube videos, enabling AI language models to access and analyze video content.1-
- AlicenseAqualityDmaintenanceMCP server that provides YouTube video data to AI agents, supporting search, metadata, comments, and transcripts without an API key.527MIT
- AlicenseNot gradedqualityCmaintenanceA comprehensive MCP server providing YouTube transcript retrieval, video search, channel browsing, playlist extraction, and upload monitoring for AI agents.10MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/BlockBenny/tubemcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server