Veo 3.1 MCP Server
Leverages the Google Files API to manage media uploads and storage, enabling token-efficient video generation by pre-uploading and reusing reference images and videos.
Integrates with Google's Veo 3.1 AI model to generate high-quality videos from text prompts and reference images, supporting features like frame interpolation and video extensions.
Facilitates the use of Google Cloud's Generative Language API and Vertex AI infrastructure for authenticated video generation, operation management, and cost-effective AI workflows.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Veo 3.1 MCP ServerGenerate an 8-second cinematic video of a cozy cabin in a snowstorm"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
π¬ Veo 3.1 MCP Server
Token-Efficient AI Video Generation with Google's Veo 3.1
π― What is This?
An MCP server for Google's Veo 3.1 - the state-of-the-art AI video generation model. Generate stunning videos from text prompts, reference images, or interpolate between first/last frames.
Key Features
β Text-to-Video - Generate videos from descriptions
β Reference Images - Up to 3 images for style guidance
β Frame Interpolation - First + last frame β coherent video
β Video Extension - Extend Veo-generated videos
β Batch Generation - Generate multiple videos with concurrency control
β Cost Estimation - Know costs before generating
β Token-Efficient - Auto-upload refs to Files API (97% token savings!)
Related MCP server: VeoMCP
π Quick Start
1. Installation
cd veo-mcp
npm install
npm run build2. Get API Key
Go to Google AI Studio
Create API key
Enable Veo 3.1 in your project (billing required)
3. Configure
cp environment.template .env
# Edit .env and add your key4. Add to Cursor
Add to ~/.cursor/mcp.json:
{
"mcpServers": {
"veo": {
"command": "node",
"args": ["C:\\Users\\woute\\Githubs\\MCP\\veo-mcp\\dist\\index.js"],
"env": {
"GEMINI_API_KEY": "your_api_key_here"
}
}
}
}Restart Cursor. Done! β
π οΈ Tools
1. start_video_generation - Generate Video
Basic text-to-video:
{
"prompt": "A serene Zen garden at sunrise, cherry blossoms falling, cinematic"
}With reference images (token-efficient!):
{
"prompt": "A futuristic cityscape at night, neon lights",
"referenceImages": [{
"source": "url",
"url": "https://example.com/style.jpg"
}],
"durationSeconds": 8,
"resolution": "1080p"
}First/last frame interpolation:
{
"prompt": "Smooth transition between these scenes",
"firstFrame": {
"source": "file_path",
"filePath": "C:\\first.jpg"
},
"lastFrame": {
"source": "file_path",
"filePath": "C:\\last.jpg"
}
}Parameters:
model-veo-3.1-generate-001(quality) orveo-3.1-fast-generate-001(speed)durationSeconds- 4, 6, or 8aspectRatio-16:9or9:16resolution-720por1080pgenerateAudio- Include synchronized audio (2x cost)seed- For reproducibilitysampleCount- Generate 1-4 videos
2. get_video_job - Check Status
{
"operationName": "operations/xyz"
}Returns status and video URLs when complete.
3. upload_image - Pre-Upload References
{
"source": "file_path",
"filePath": "C:\\style-ref.jpg"
}Returns fileUri valid for 48 hours. Reuse across multiple generations!
4. extend_video - Extend Videos
{
"videoFileUri": "files/abc123",
"additionalSeconds": 7,
"prompt": "Continue with the character walking into the sunset"
}5. start_batch_video_generation - Batch Generate
{
"jobs": [
{"key": "scene1", "request": {"prompt": "..."}},
{"key": "scene2", "request": {"prompt": "..."}}
],
"concurrency": 3
}6. estimate_veo_cost - Cost Estimation
{
"model": "veo-3.1-fast-generate-001",
"durationSeconds": 8,
"sampleCount": 1,
"generateAudio": false
}Returns estimated cost in USD.
π° Pricing
Model | Video Only | Video + Audio |
veo-3.1-generate-001 (quality) | $0.20/sec | $0.40/sec |
veo-3.1-fast-generate-001 (speed) | $0.10/sec | $0.15/sec |
Example Costs:
8s video (fast, no audio): $0.80
8s video (quality, with audio): $3.20
4s video (fast, no audio): $0.40
π Limits & Constraints
Parameter | Limit |
Duration | 4, 6, or 8 seconds |
Reference images | 0-3 images |
Sample count | 1-4 videos |
Resolutions | 720p, 1080p |
Aspect ratios | 16:9, 9:16 |
Rate limit | ~50 requests/min |
π‘ Usage Examples
Simple Text-to-Video
Generate an 8-second video of a peaceful forest scene with morning mistWith Style Reference
Create a video of a tech startup office, using this image for style: C:\ref.jpgFrame Interpolation
Generate a smooth transition between first.jpg and last.jpg, 8 seconds, cinematic camera movementBatch Generation
Generate 5 different video variations of a product showcase with different anglesπ How Token Efficiency Works
β Naive Approach (Base64)
{
"referenceImages": [{
"base64": "iVBORw0KGgo..." // 500KB β ~50,000 tokens!
}]
}Cost: Massive token usage per call
β Token-Efficient (This MCP)
{
"referenceImages": [{
"source": "url",
"url": "https://example.com/ref.jpg" // ~20 tokens
}]
}What Happens:
Server downloads image (no tokens)
Computes SHA-256 hash
Checks cache (48h validity)
Uploads to Files API if needed (~1s)
Uses short
files/abc123URI (~5 tokens)
Savings: 97%+ fewer tokens! π
β±οΈ Generation Times
Configuration | Typical Time |
4s, 720p, no audio | 30-60 sec |
8s, 1080p, no audio | 60-120 sec |
8s, 1080p, with audio | 90-150 sec |
With references | +10-30 sec |
Frame interpolation | +20-40 sec |
Note: Times vary based on prompt complexity and server load.
π¨ Best Practices
1. Start Small, Scale Up
Step 1: Generate 1 video at 720p
Step 2: If good, regenerate at 1080p
Step 3: Use batch for variations2. Use Fast Model for Testing
{
"model": "veo-3.1-fast-generate-001", // Testing
"resolution": "720p"
}Switch to quality model for final:
{
"model": "veo-3.1-generate-001", // Final
"resolution": "1080p"
}3. Pre-Upload Frequently Used References
// Step 1: Upload once
upload_image {"source": "file_path", "filePath": "brand-style.jpg"}
// Returns: files/xyz123
// Step 2: Reuse many times
{
"referenceImages": [{"source": "file_uri", "fileUri": "files/xyz123"}]
}4. Leverage Batch for Variations
{
"jobs": [
{"key": "v1", "request": {"prompt": "Scene 1...", "seed": 1}},
{"key": "v2", "request": {"prompt": "Scene 1...", "seed": 2}},
{"key": "v3", "request": {"prompt": "Scene 1...", "seed": 3}}
]
}5. Monitor Costs
Always estimate before large batches:
estimate_veo_cost {
"model": "veo-3.1-fast-generate-001",
"durationSeconds": 8,
"sampleCount": 10
}
// Returns: $8.00 estimate㪠Async Operation Flow
Veo uses async long-running operations:
1. start_video_generation
β Returns operationName immediately
2. get_video_job (poll every 10-30s)
β Returns {done: false, status: "RUNNING"}
3. get_video_job (after 60-120s)
β Returns {done: true, videos: [{videoUri: "..."}]}
4. Download video from videoUriTip: Don't poll too frequently (< 10s intervals).
π Troubleshooting
"API not enabled" (403)
Go to Google Cloud Console
Enable "Generative Language API"
Enable billing
Wait 5-10 minutes for propagation
"Rate limit exceeded"
Veo allows ~50 requests/min
Use batch tool with
concurrency: 3Add delays between requests
"Invalid aspect ratio with references"
9:16 may not work with reference images
Use 16:9 for reference mode
Check Veo 3.1 docs for updates
"Video extension failed"
Only Veo-generated videos can be extended
Cannot extend arbitrary MP4s
Input must be from previous Veo job
Long generation times
1080p takes longer than 720p
Audio generation adds time
Reference images add processing
Frame interpolation is slowest
π Resources
π― Status: Production Ready β
β All 6 tools implemented
β Token-efficient file handling
β Async operation support
β Batch generation with concurrency control
β Cost estimation
β Comprehensive validation
β Error handling
β Full documentation
Ready to generate amazing videos! π
Built with π¬ for AI video generation
Available Tools
6 toolsestimate_veo_costB
Estimate the cost in USD for a video generation request before starting it. Helps plan budgets and batch sizes.
| Name | Required | Description | Default |
|---|---|---|---|
| model | Yes | Model to use | |
| durationSeconds | Yes | Video duration | |
| sampleCount | No | Number of videos (default: 1) | |
| generateAudio | No | Whether to generate audio (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. While it mentions the tool's purpose and timing, it doesn't describe what the estimation process entails, whether it requires authentication, if it has rate limits, what the return format looks like, or any error conditions. For a cost estimation tool with zero annotation coverage, this leaves significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with just two sentences that each earn their place. The first sentence states the core purpose and timing, while the second explains the value proposition. There's zero wasted text, and the information is front-loaded appropriately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a cost estimation tool with no annotations and no output schema, the description is insufficiently complete. It doesn't explain what the estimation returns (just 'cost in USD' without format details), doesn't mention whether this is a real API call or local calculation, and provides minimal behavioral context. The description should do more to compensate for the lack of structured metadata.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema. However, it does provide context about why parameters matter ('Helps plan budgets and batch sizes'), which adds marginal value. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: estimating cost in USD for video generation requests. It specifies the action ('estimate cost'), resource ('video generation request'), and context ('before starting it'). However, it doesn't explicitly differentiate from sibling tools like 'start_video_generation' beyond the 'before starting' timing hint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides implied usage context: 'Helps plan budgets and batch sizes' and 'before starting it' suggests this should be used prior to actual generation tools. However, it doesn't explicitly state when to use this versus alternatives or mention specific prerequisites beyond the timing suggestion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
extend_videoB
Extend a Veo-generated video by additional seconds. Input video must be from a previous Veo generation (not an arbitrary video).
| Name | Required | Description | Default |
|---|---|---|---|
| videoFileUri | Yes | FileUri of the Veo-generated video to extend | |
| additionalSeconds | Yes | Number of seconds to add (typically 7 for 8s extension) | |
| prompt | No | Optional: Continuation prompt describing what should happen next | |
| model | No | Model to use (default: fast) | |
| seed | No | Optional seed for reproducibility |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions the input constraint (Veo-generated video only) but lacks details on permissions, rate limits, whether the extension is reversible, or what the output looks like (e.g., file format, processing time). For a mutation tool with zero annotation coverage, this is a significant gap in behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero waste, front-loaded with the core purpose and followed by a key constraint. Every sentence earns its place by adding essential information without redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 5 parameters, no annotations, and no output schema, the description is adequate but incomplete. It covers the purpose and input constraint but lacks behavioral details (e.g., mutation effects, error handling) and output information, which are important for a video extension tool with multiple parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond the schema by hinting at the input constraint ('from a previous Veo generation'), but doesn't provide additional syntax, format details, or usage examples for parameters. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('extend') and resource ('Veo-generated video'), specifying it adds seconds to a video. It distinguishes from arbitrary video processing by noting the input must be from previous Veo generation, but doesn't explicitly differentiate from sibling tools like 'start_video_generation' or 'get_video_job'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by stating the input video must be from a previous Veo generation, which provides some context. However, it doesn't specify when to use this tool versus alternatives like 'start_video_generation' for new videos or 'get_video_job' for checking status, nor does it mention prerequisites or exclusions beyond the input requirement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_video_jobA
Check the status of a video generation job. Returns status and video URLs when complete. Videos are returned as download URLs - you can save them locally or share the links.
| Name | Required | Description | Default |
|---|---|---|---|
| operationName | Yes | Operation name from start_video_generation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It adds useful context by describing the return format (status and video URLs as download URLs) and usage options (save locally or share links), but it does not cover potential errors, rate limits, or authentication needs, leaving gaps for a mutation-free tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is appropriately sized and front-loaded, with two sentences that efficiently convey the tool's purpose, return values, and usage without wasted words. Each sentence adds distinct value, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (one parameter, no output schema, no annotations), the description is mostly complete, covering purpose, return format, and usage. However, it lacks details on error handling or response structure, which could be helpful for an AI agent, though not critical for this simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the single parameter (operationName from start_video_generation). The description does not add any parameter-specific details beyond what the schema provides, such as format examples or constraints, meeting the baseline for high schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Check') and resource ('video generation job'), distinguishing it from siblings like start_video_generation (which initiates jobs) and extend_video (which modifies videos). It explicitly mentions what it returns (status and video URLs), making the function unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for usage by specifying that it checks job status and returns results when complete, implying it should be used after starting a job. However, it does not explicitly state when not to use it or name alternatives, such as using start_video_generation for new jobs instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_batch_video_generationA
Start multiple video generation jobs with controlled concurrency. Returns operation names for all jobs. Use this to generate multiple videos efficiently while respecting rate limits.
| Name | Required | Description | Default |
|---|---|---|---|
| jobs | Yes | Array of video generation requests | |
| concurrency | No | Max concurrent requests (default: 3, recommend <= 5 to avoid rate limits) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses key behavioral traits: it starts multiple jobs with concurrency control, returns operation names, and respects rate limits. However, it lacks details on error handling, timeouts, or what 'operation names' entail, leaving gaps for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose in the first sentence, followed by usage guidance. Every sentence earns its place by adding value (e.g., return values and rate limit context), with no wasted words, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a batch mutation tool with no annotations and no output schema, the description is moderately complete. It covers the purpose and high-level behavior but lacks details on output format, error scenarios, or prerequisites, which are important for such an operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description adds minimal value beyond the schema, mentioning 'controlled concurrency' which aligns with the 'concurrency' parameter but doesn't provide additional syntax or format details. Baseline 3 is appropriate as the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Start multiple video generation jobs') and resource ('video generation jobs'), distinguishing it from siblings like 'start_video_generation' (singular) and 'get_video_job' (retrieval). It specifies the concurrency control aspect, making the purpose specific and differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool ('to generate multiple videos efficiently while respecting rate limits'), but it does not explicitly state when not to use it or name alternatives. For example, it doesn't clarify whether to use this over 'start_video_generation' for single jobs or other siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_video_generationA
Start a Veo 3.1 video generation job. This returns an operation ID immediately - use get_video_job to poll for completion. Supports text-to-video, reference images (up to 3), and first/last frame interpolation.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | Text description of the video to generate | |
| model | No | Model to use: veo-3.1-generate-preview (quality, $0.75/sec) or veo-3.1-fast-generate-preview (speed, $0.10/sec). Default: fast | |
| durationSeconds | No | Video duration: 4, 6, or 8 seconds (default: 8) | |
| aspectRatio | No | Aspect ratio (default: 16:9). Note: 9:16 may not work with reference images. | |
| resolution | No | Video resolution (default: 1080p) | |
| seed | No | Optional seed for reproducible generation | |
| sampleCount | No | Number of videos to generate (1-4, default: 1) | |
| generateAudio | No | Whether to generate synchronized audio (default: false). Costs 2x more. | |
| referenceImages | No | Up to 3 reference images for visual guidance. Each can be URL, file path, or fileUri. | |
| firstFrame | No | First frame for interpolation (must also provide lastFrame) | |
| lastFrame | No | Last frame for interpolation (must also provide firstFrame) | |
| negativePrompt | No | Optional: Things to avoid in the video | |
| resizeMode | No | How to fit reference images (default: pad) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the asynchronous nature ('returns an operation ID immediately') and polling requirement, which is crucial. However, it lacks information about costs, rate limits, authentication needs, or error handling, which are important for a complex video generation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise with three sentences that each earn their place: the core purpose, the polling requirement, and the key capabilities. It's front-loaded with the most critical information and wastes no words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (13 parameters, nested objects) and lack of both annotations and output schema, the description is somewhat incomplete. While it covers the asynchronous behavior and key features, it doesn't address costs, permissions, or what the operation ID represents. For such a rich tool, more contextual guidance would be helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 100%, so the schema already documents all 13 parameters thoroughly. The description adds minimal value by mentioning 'reference images (up to 3)' and 'first/last frame interpolation', which are already covered in the schema. This meets the baseline of 3 when schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Start a Veo 3.1 video generation job') and resource ('video generation job'), distinguishing it from siblings like 'extend_video' or 'start_batch_video_generation'. It also mentions the immediate return type ('operation ID') and key capabilities ('text-to-video, reference images, first/last frame interpolation').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use this tool by stating it 'returns an operation ID immediately - use get_video_job to poll for completion', which implicitly guides the agent to follow up with the sibling tool. However, it doesn't explicitly mention when NOT to use it or alternatives like 'start_batch_video_generation' for batch jobs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
upload_imageA
Upload an image to Google Files API for use as reference, first frame, or last frame in video generation. Returns a fileUri that can be reused for 48 hours. This is the most token-efficient way to pass images to video generation.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | Source type: 'url' to download from web, 'file_path' for local file | |
| url | No | URL to download (if source='url') | |
| filePath | No | Local file path (if source='file_path') | |
| displayName | No | Optional display name for the uploaded file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behaviors: the tool returns a 'fileUri that can be reused for 48 hours' (temporal constraint) and is 'the most token-efficient way to pass images' (performance characteristic). However, it doesn't mention authentication requirements, rate limits, or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is perfectly concise with three sentences that each serve distinct purposes: stating the core function, describing the return value and constraints, and providing performance context. Every sentence earns its place with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description does well by explaining the return value ('fileUri') and its 48-hour validity. However, it could be more complete by mentioning what happens with invalid inputs, file size limits, or supported image formats given the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema description coverage, the schema already documents all 4 parameters thoroughly. The description adds no additional parameter information beyond what's in the schema, so it meets the baseline expectation but doesn't provide extra value regarding parameter usage or semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('upload an image'), target resource ('Google Files API'), and primary use case ('for use as reference, first frame, or last frame in video generation'). It distinguishes this tool from sibling video generation tools by focusing on image preparation rather than video creation or management.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context about when to use this tool ('for use as reference, first frame, or last frame in video generation') and mentions it's 'the most token-efficient way to pass images to video generation.' However, it doesn't explicitly state when NOT to use it or name specific alternatives among the sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
6 tool updates
v0.1.0- First observed
estimate_veo_cost - First observed
extend_video - First observed
get_video_job - First observed
start_batch_video_generation - First observed
start_video_generation - First observed
upload_image
TDQS
Each tool has a clearly distinct purpose with no overlap: cost estimation, video extension, job status checking, batch generation, single generation, and image upload. The descriptions reinforce distinct workflows (e.g., 'extend_video' specifically requires Veo-generated input, while 'upload_image' handles image preparation).
All tools follow a consistent verb_noun pattern with snake_case (e.g., 'estimate_veo_cost', 'extend_video', 'get_video_job'). The naming is predictable and readable, using clear verbs like 'estimate', 'extend', 'get', 'start', and 'upload' paired with relevant nouns.
Six tools are well-scoped for a video generation server, covering core operations from initiation to status checking and cost management. This count avoids bloat while ensuring each tool earns its place, such as separate tools for single and batch generation to handle different use cases efficiently.
The toolset provides complete coverage for the video generation domain: starting jobs (single and batch), checking status, extending videos, uploading images for reference, and estimating costs. There are no obvious gaps; agents can handle the full lifecycle from planning to retrieval without dead ends.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Create images and videos from prompts, with options for image mixing, reference images, and start/β¦
AI image, video & music generation. Flux, Veo 3.1, Suno V5. Free tier included.
AI image, video, voice and music generation over MCP, routed to Veo 3.1, Seedance 2.0 and more.
Best Image and video generation: 20+ models (Kling, Seedance, Veo, NB, FLUX.2), OAuth, pay-per-use.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables video generation from text prompts or images using Google's Veo 3 API. Supports multiple models, audio generation, and various aspect ratios for creating high-quality videos.2MIT
- AlicenseAqualityCmaintenanceGoogle Veo AI video generation with text-to-video, image-to-video, multi-image fusion, 1080p upscaling, and multiple quality/speed models.83MIT
- AlicenseAqualityCmaintenanceMCP server for Google Veo 3.1 video generation. Supports text/video/image-based generation, extension, and interpolation with cost estimation and batch processing.437MIT
- AlicenseAqualityDmaintenanceEnables image generation, editing, and analysis using Google's Gemini 2.5 Flash and Gemini 3 Pro models, with support for batch processing, style templates, and high-resolution output.87581MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/waimakers/veo-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server