Skip to main content
Glama

🎬 Veo 3.1 MCP Server

Token-Efficient AI Video Generation with Google's Veo 3.1

🎯 What is This?

An MCP server for Google's Veo 3.1 - the state-of-the-art AI video generation model. Generate stunning videos from text prompts, reference images, or interpolate between first/last frames.

Key Features

  • βœ… Text-to-Video - Generate videos from descriptions

  • βœ… Reference Images - Up to 3 images for style guidance

  • βœ… Frame Interpolation - First + last frame β†’ coherent video

  • βœ… Video Extension - Extend Veo-generated videos

  • βœ… Batch Generation - Generate multiple videos with concurrency control

  • βœ… Cost Estimation - Know costs before generating

  • βœ… Token-Efficient - Auto-upload refs to Files API (97% token savings!)


Related MCP server: VeoMCP

πŸš€ Quick Start

1. Installation

cd veo-mcp
npm install
npm run build

2. Get API Key

  1. Go to Google AI Studio

  2. Create API key

  3. Enable Veo 3.1 in your project (billing required)

3. Configure

cp environment.template .env
# Edit .env and add your key

4. Add to Cursor

Add to ~/.cursor/mcp.json:

{
  "mcpServers": {
    "veo": {
      "command": "node",
      "args": ["C:\\Users\\woute\\Githubs\\MCP\\veo-mcp\\dist\\index.js"],
      "env": {
        "GEMINI_API_KEY": "your_api_key_here"
      }
    }
  }
}

Restart Cursor. Done! βœ…


πŸ› οΈ Tools

1. start_video_generation - Generate Video

Basic text-to-video:

{
  "prompt": "A serene Zen garden at sunrise, cherry blossoms falling, cinematic"
}

With reference images (token-efficient!):

{
  "prompt": "A futuristic cityscape at night, neon lights",
  "referenceImages": [{
    "source": "url",
    "url": "https://example.com/style.jpg"
  }],
  "durationSeconds": 8,
  "resolution": "1080p"
}

First/last frame interpolation:

{
  "prompt": "Smooth transition between these scenes",
  "firstFrame": {
    "source": "file_path",
    "filePath": "C:\\first.jpg"
  },
  "lastFrame": {
    "source": "file_path",
    "filePath": "C:\\last.jpg"
  }
}

Parameters:

  • model - veo-3.1-generate-001 (quality) or veo-3.1-fast-generate-001 (speed)

  • durationSeconds - 4, 6, or 8

  • aspectRatio - 16:9 or 9:16

  • resolution - 720p or 1080p

  • generateAudio - Include synchronized audio (2x cost)

  • seed - For reproducibility

  • sampleCount - Generate 1-4 videos

2. get_video_job - Check Status

{
  "operationName": "operations/xyz"
}

Returns status and video URLs when complete.

3. upload_image - Pre-Upload References

{
  "source": "file_path",
  "filePath": "C:\\style-ref.jpg"
}

Returns fileUri valid for 48 hours. Reuse across multiple generations!

4. extend_video - Extend Videos

{
  "videoFileUri": "files/abc123",
  "additionalSeconds": 7,
  "prompt": "Continue with the character walking into the sunset"
}

5. start_batch_video_generation - Batch Generate

{
  "jobs": [
    {"key": "scene1", "request": {"prompt": "..."}},
    {"key": "scene2", "request": {"prompt": "..."}}
  ],
  "concurrency": 3
}

6. estimate_veo_cost - Cost Estimation

{
  "model": "veo-3.1-fast-generate-001",
  "durationSeconds": 8,
  "sampleCount": 1,
  "generateAudio": false
}

Returns estimated cost in USD.


πŸ’° Pricing

Model

Video Only

Video + Audio

veo-3.1-generate-001 (quality)

$0.20/sec

$0.40/sec

veo-3.1-fast-generate-001 (speed)

$0.10/sec

$0.15/sec

Example Costs:

  • 8s video (fast, no audio): $0.80

  • 8s video (quality, with audio): $3.20

  • 4s video (fast, no audio): $0.40


πŸ“Š Limits & Constraints

Parameter

Limit

Duration

4, 6, or 8 seconds

Reference images

0-3 images

Sample count

1-4 videos

Resolutions

720p, 1080p

Aspect ratios

16:9, 9:16

Rate limit

~50 requests/min


πŸ’‘ Usage Examples

Simple Text-to-Video

Generate an 8-second video of a peaceful forest scene with morning mist

With Style Reference

Create a video of a tech startup office, using this image for style: C:\ref.jpg

Frame Interpolation

Generate a smooth transition between first.jpg and last.jpg, 8 seconds, cinematic camera movement

Batch Generation

Generate 5 different video variations of a product showcase with different angles

πŸ” How Token Efficiency Works

❌ Naive Approach (Base64)

{
  "referenceImages": [{
    "base64": "iVBORw0KGgo..." // 500KB β†’ ~50,000 tokens!
  }]
}

Cost: Massive token usage per call

βœ… Token-Efficient (This MCP)

{
  "referenceImages": [{
    "source": "url",
    "url": "https://example.com/ref.jpg" // ~20 tokens
  }]
}

What Happens:

  1. Server downloads image (no tokens)

  2. Computes SHA-256 hash

  3. Checks cache (48h validity)

  4. Uploads to Files API if needed (~1s)

  5. Uses short files/abc123 URI (~5 tokens)

Savings: 97%+ fewer tokens! πŸŽ‰


⏱️ Generation Times

Configuration

Typical Time

4s, 720p, no audio

30-60 sec

8s, 1080p, no audio

60-120 sec

8s, 1080p, with audio

90-150 sec

With references

+10-30 sec

Frame interpolation

+20-40 sec

Note: Times vary based on prompt complexity and server load.


🎨 Best Practices

1. Start Small, Scale Up

Step 1: Generate 1 video at 720p
Step 2: If good, regenerate at 1080p
Step 3: Use batch for variations

2. Use Fast Model for Testing

{
  "model": "veo-3.1-fast-generate-001",  // Testing
  "resolution": "720p"
}

Switch to quality model for final:

{
  "model": "veo-3.1-generate-001",  // Final
  "resolution": "1080p"
}

3. Pre-Upload Frequently Used References

// Step 1: Upload once
upload_image {"source": "file_path", "filePath": "brand-style.jpg"}
// Returns: files/xyz123

// Step 2: Reuse many times
{
  "referenceImages": [{"source": "file_uri", "fileUri": "files/xyz123"}]
}

4. Leverage Batch for Variations

{
  "jobs": [
    {"key": "v1", "request": {"prompt": "Scene 1...", "seed": 1}},
    {"key": "v2", "request": {"prompt": "Scene 1...", "seed": 2}},
    {"key": "v3", "request": {"prompt": "Scene 1...", "seed": 3}}
  ]
}

5. Monitor Costs

Always estimate before large batches:

estimate_veo_cost {
  "model": "veo-3.1-fast-generate-001",
  "durationSeconds": 8,
  "sampleCount": 10
}
// Returns: $8.00 estimate

🎬 Async Operation Flow

Veo uses async long-running operations:

1. start_video_generation
   ↓ Returns operationName immediately
   
2. get_video_job (poll every 10-30s)
   ↓ Returns {done: false, status: "RUNNING"}
   
3. get_video_job (after 60-120s)
   ↓ Returns {done: true, videos: [{videoUri: "..."}]}
   
4. Download video from videoUri

Tip: Don't poll too frequently (< 10s intervals).


πŸ†˜ Troubleshooting

"API not enabled" (403)

  1. Go to Google Cloud Console

  2. Enable "Generative Language API"

  3. Enable billing

  4. Wait 5-10 minutes for propagation

"Rate limit exceeded"

  • Veo allows ~50 requests/min

  • Use batch tool with concurrency: 3

  • Add delays between requests

"Invalid aspect ratio with references"

  • 9:16 may not work with reference images

  • Use 16:9 for reference mode

  • Check Veo 3.1 docs for updates

"Video extension failed"

  • Only Veo-generated videos can be extended

  • Cannot extend arbitrary MP4s

  • Input must be from previous Veo job

Long generation times

  • 1080p takes longer than 720p

  • Audio generation adds time

  • Reference images add processing

  • Frame interpolation is slowest


πŸ“š Resources


🎯 Status: Production Ready βœ…

  • βœ… All 6 tools implemented

  • βœ… Token-efficient file handling

  • βœ… Async operation support

  • βœ… Batch generation with concurrency control

  • βœ… Cost estimation

  • βœ… Comprehensive validation

  • βœ… Error handling

  • βœ… Full documentation

Ready to generate amazing videos! πŸš€


Built with 🎬 for AI video generation

Available Tools

6 tools
estimate_veo_costB

Estimate the cost in USD for a video generation request before starting it. Helps plan budgets and batch sizes.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelYesModel to use
durationSecondsYesVideo duration
sampleCountNoNumber of videos (default: 1)
generateAudioNoWhether to generate audio (default: false)

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden for behavioral disclosure. While it mentions the tool's purpose and timing, it doesn't describe what the estimation process entails, whether it requires authentication, if it has rate limits, what the return format looks like, or any error conditions. For a cost estimation tool with zero annotation coverage, this leaves significant behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise with just two sentences that each earn their place. The first sentence states the core purpose and timing, while the second explains the value proposition. There's zero wasted text, and the information is front-loaded appropriately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a cost estimation tool with no annotations and no output schema, the description is insufficiently complete. It doesn't explain what the estimation returns (just 'cost in USD' without format details), doesn't mention whether this is a real API call or local calculation, and provides minimal behavioral context. The description should do more to compensate for the lack of structured metadata.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description doesn't add any parameter-specific information beyond what's in the schema. However, it does provide context about why parameters matter ('Helps plan budgets and batch sizes'), which adds marginal value. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: estimating cost in USD for video generation requests. It specifies the action ('estimate cost'), resource ('video generation request'), and context ('before starting it'). However, it doesn't explicitly differentiate from sibling tools like 'start_video_generation' beyond the 'before starting' timing hint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides implied usage context: 'Helps plan budgets and batch sizes' and 'before starting it' suggests this should be used prior to actual generation tools. However, it doesn't explicitly state when to use this versus alternatives or mention specific prerequisites beyond the timing suggestion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

extend_videoB

Extend a Veo-generated video by additional seconds. Input video must be from a previous Veo generation (not an arbitrary video).

ParametersJSON Schema
NameRequiredDescriptionDefault
videoFileUriYesFileUri of the Veo-generated video to extend
additionalSecondsYesNumber of seconds to add (typically 7 for 8s extension)
promptNoOptional: Continuation prompt describing what should happen next
modelNoModel to use (default: fast)
seedNoOptional seed for reproducibility

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions the input constraint (Veo-generated video only) but lacks details on permissions, rate limits, whether the extension is reversible, or what the output looks like (e.g., file format, processing time). For a mutation tool with zero annotation coverage, this is a significant gap in behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with zero waste, front-loaded with the core purpose and followed by a key constraint. Every sentence earns its place by adding essential information without redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters, no annotations, and no output schema, the description is adequate but incomplete. It covers the purpose and input constraint but lacks behavioral details (e.g., mutation effects, error handling) and output information, which are important for a video extension tool with multiple parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds minimal value beyond the schema by hinting at the input constraint ('from a previous Veo generation'), but doesn't provide additional syntax, format details, or usage examples for parameters. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('extend') and resource ('Veo-generated video'), specifying it adds seconds to a video. It distinguishes from arbitrary video processing by noting the input must be from previous Veo generation, but doesn't explicitly differentiate from sibling tools like 'start_video_generation' or 'get_video_job'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by stating the input video must be from a previous Veo generation, which provides some context. However, it doesn't specify when to use this tool versus alternatives like 'start_video_generation' for new videos or 'get_video_job' for checking status, nor does it mention prerequisites or exclusions beyond the input requirement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_video_jobA

Check the status of a video generation job. Returns status and video URLs when complete. Videos are returned as download URLs - you can save them locally or share the links.

ParametersJSON Schema
NameRequiredDescriptionDefault
operationNameYesOperation name from start_video_generation

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds useful context by describing the return format (status and video URLs as download URLs) and usage options (save locally or share links), but it does not cover potential errors, rate limits, or authentication needs, leaving gaps for a mutation-free tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded, with two sentences that efficiently convey the tool's purpose, return values, and usage without wasted words. Each sentence adds distinct value, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's low complexity (one parameter, no output schema, no annotations), the description is mostly complete, covering purpose, return format, and usage. However, it lacks details on error handling or response structure, which could be helpful for an AI agent, though not critical for this simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the single parameter (operationName from start_video_generation). The description does not add any parameter-specific details beyond what the schema provides, such as format examples or constraints, meeting the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Check') and resource ('video generation job'), distinguishing it from siblings like start_video_generation (which initiates jobs) and extend_video (which modifies videos). It explicitly mentions what it returns (status and video URLs), making the function unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for usage by specifying that it checks job status and returns results when complete, implying it should be used after starting a job. However, it does not explicitly state when not to use it or name alternatives, such as using start_video_generation for new jobs instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_batch_video_generationA

Start multiple video generation jobs with controlled concurrency. Returns operation names for all jobs. Use this to generate multiple videos efficiently while respecting rate limits.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobsYesArray of video generation requests
concurrencyNoMax concurrent requests (default: 3, recommend <= 5 to avoid rate limits)

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses key behavioral traits: it starts multiple jobs with concurrency control, returns operation names, and respects rate limits. However, it lacks details on error handling, timeouts, or what 'operation names' entail, leaving gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose in the first sentence, followed by usage guidance. Every sentence earns its place by adding value (e.g., return values and rate limit context), with no wasted words, making it highly efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity of a batch mutation tool with no annotations and no output schema, the description is moderately complete. It covers the purpose and high-level behavior but lacks details on output format, error scenarios, or prerequisites, which are important for such an operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description adds minimal value beyond the schema, mentioning 'controlled concurrency' which aligns with the 'concurrency' parameter but doesn't provide additional syntax or format details. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Start multiple video generation jobs') and resource ('video generation jobs'), distinguishing it from siblings like 'start_video_generation' (singular) and 'get_video_job' (retrieval). It specifies the concurrency control aspect, making the purpose specific and differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool ('to generate multiple videos efficiently while respecting rate limits'), but it does not explicitly state when not to use it or name alternatives. For example, it doesn't clarify whether to use this over 'start_video_generation' for single jobs or other siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

start_video_generationA

Start a Veo 3.1 video generation job. This returns an operation ID immediately - use get_video_job to poll for completion. Supports text-to-video, reference images (up to 3), and first/last frame interpolation.

ParametersJSON Schema
NameRequiredDescriptionDefault
promptYesText description of the video to generate
modelNoModel to use: veo-3.1-generate-preview (quality, $0.75/sec) or veo-3.1-fast-generate-preview (speed, $0.10/sec). Default: fast
durationSecondsNoVideo duration: 4, 6, or 8 seconds (default: 8)
aspectRatioNoAspect ratio (default: 16:9). Note: 9:16 may not work with reference images.
resolutionNoVideo resolution (default: 1080p)
seedNoOptional seed for reproducible generation
sampleCountNoNumber of videos to generate (1-4, default: 1)
generateAudioNoWhether to generate synchronized audio (default: false). Costs 2x more.
referenceImagesNoUp to 3 reference images for visual guidance. Each can be URL, file path, or fileUri.
firstFrameNoFirst frame for interpolation (must also provide lastFrame)
lastFrameNoLast frame for interpolation (must also provide firstFrame)
negativePromptNoOptional: Things to avoid in the video
resizeModeNoHow to fit reference images (default: pad)

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes the asynchronous nature ('returns an operation ID immediately') and polling requirement, which is crucial. However, it lacks information about costs, rate limits, authentication needs, or error handling, which are important for a complex video generation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is perfectly concise with three sentences that each earn their place: the core purpose, the polling requirement, and the key capabilities. It's front-loaded with the most critical information and wastes no words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complexity (13 parameters, nested objects) and lack of both annotations and output schema, the description is somewhat incomplete. While it covers the asynchronous behavior and key features, it doesn't address costs, permissions, or what the operation ID represents. For such a rich tool, more contextual guidance would be helpful.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the schema already documents all 13 parameters thoroughly. The description adds minimal value by mentioning 'reference images (up to 3)' and 'first/last frame interpolation', which are already covered in the schema. This meets the baseline of 3 when schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Start a Veo 3.1 video generation job') and resource ('video generation job'), distinguishing it from siblings like 'extend_video' or 'start_batch_video_generation'. It also mentions the immediate return type ('operation ID') and key capabilities ('text-to-video, reference images, first/last frame interpolation').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on when to use this tool by stating it 'returns an operation ID immediately - use get_video_job to poll for completion', which implicitly guides the agent to follow up with the sibling tool. However, it doesn't explicitly mention when NOT to use it or alternatives like 'start_batch_video_generation' for batch jobs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload_imageA

Upload an image to Google Files API for use as reference, first frame, or last frame in video generation. Returns a fileUri that can be reused for 48 hours. This is the most token-efficient way to pass images to video generation.

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYesSource type: 'url' to download from web, 'file_path' for local file
urlNoURL to download (if source='url')
filePathNoLocal file path (if source='file_path')
displayNameNoOptional display name for the uploaded file

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It effectively describes key behaviors: the tool returns a 'fileUri that can be reused for 48 hours' (temporal constraint) and is 'the most token-efficient way to pass images' (performance characteristic). However, it doesn't mention authentication requirements, rate limits, or error conditions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is perfectly concise with three sentences that each serve distinct purposes: stating the core function, describing the return value and constraints, and providing performance context. Every sentence earns its place with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description does well by explaining the return value ('fileUri') and its 48-hour validity. However, it could be more complete by mentioning what happens with invalid inputs, file size limits, or supported image formats given the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 100% schema description coverage, the schema already documents all 4 parameters thoroughly. The description adds no additional parameter information beyond what's in the schema, so it meets the baseline expectation but doesn't provide extra value regarding parameter usage or semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('upload an image'), target resource ('Google Files API'), and primary use case ('for use as reference, first frame, or last frame in video generation'). It distinguishes this tool from sibling video generation tools by focusing on image preparation rather than video creation or management.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context about when to use this tool ('for use as reference, first frame, or last frame in video generation') and mentions it's 'the most token-efficient way to pass images to video generation.' However, it doesn't explicitly state when NOT to use it or name specific alternatives among the sibling tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 6 tool updatesv0.1.0
    • First observedestimate_veo_cost
    • First observedextend_video
    • First observedget_video_job
    • First observedstart_batch_video_generation
    • First observedstart_video_generation
    • First observedupload_image

TDQS

A4/5.0
Disambiguation5/5

Each tool has a clearly distinct purpose with no overlap: cost estimation, video extension, job status checking, batch generation, single generation, and image upload. The descriptions reinforce distinct workflows (e.g., 'extend_video' specifically requires Veo-generated input, while 'upload_image' handles image preparation).

Naming Consistency5/5

All tools follow a consistent verb_noun pattern with snake_case (e.g., 'estimate_veo_cost', 'extend_video', 'get_video_job'). The naming is predictable and readable, using clear verbs like 'estimate', 'extend', 'get', 'start', and 'upload' paired with relevant nouns.

Tool Count5/5

Six tools are well-scoped for a video generation server, covering core operations from initiation to status checking and cost management. This count avoids bloat while ensuring each tool earns its place, such as separate tools for single and batch generation to handle different use cases efficiently.

Completeness5/5

The toolset provides complete coverage for the video generation domain: starting jobs (single and batch), checking status, extending videos, uploading images for reference, and estimating costs. There are no obvious gaps; agents can handle the full lifecycle from planning to retrieval without dead ends.

Maintenance

ActivityInactive
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/waimakers/veo-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server