Skip to main content
Glama

Generate Video

clipform_generate_video
Destructive

Generate a video from images, video clips, or both, synced to an audio track. Use this for narrated question backgrounds, topic visualisations, or any form node that benefits from video. Combine with clipform_generate_tts for narrated audio and clipform_search_media for royalty-free images. Creates 9:16 (720x1280) with Ken Burns pan/zoom effects and transitions. Returns a public URL when complete.

Items: type "image" (Ken Burns motion) or "video" (cover-cropped, muted by default). Duration matches audio_url or set duration_seconds explicitly.

For multi-question builds, pass wait: false on every render: each call returns a job ID immediately, so all renders run in parallel - then collect URLs with clipform_check_render. Sequential waiting renders take 15-120 seconds EACH.

Choosing a render tool: for a recognisable form/quiz beat (guess-the-city, this-or-that, mystery reveal, multiple choice, photo montage...) reach for a video template first (clipform_list_video_templates + clipform_render_video_template) - it is a one-call recipe. Use clipform_generate_video for a narrated or audio-synced media montage (images/clips timed to a voice track). Use clipform_render_composition only when neither fits and you need a custom layer stack. Montage disambiguation: choose clipform_generate_video when the montage is narrated or synced to an audio track; choose the slideshow video template when it is silent (motion + transitions only, no voice-over). A render for a form node is not done until it is attached to that node. Pass node_id (and form_id) so the completed render attaches itself automatically - do not poll clipform_check_render to completion or manually chain clipform_upload_media_asset + clipform_attach_node_media; fire the render and move on.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
waitNotrue (default) blocks until the video is ready and returns its URL. false returns a job ID immediately - fire all renders first, then poll clipform_check_render. Use false whenever rendering more than one video.
itemsYesMedia items (images, video clips, or a mix)
contextYesDescribe the user's underlying goal in one sentence - not the tool you're calling.
duotoneNoTwo-tone editorial recolour on image items - desaturates then maps to a shadow->highlight palette. Pair with a halftone texture for a screen-print poster look.
form_idNoThe form UUID (required when node_id is set).
node_idNoForm node to attach this render to automatically once it completes - skips the manual clipform_upload_media_asset + clipform_attach_node_media steps and republishes the form if it's currently live. Requires form_id.
textureNoPrint-style pattern overlay on image items - makes stock imagery read as designed (screen-print dither look)
captionsNoWord-level captions from clipform_generate_tts - carried onto the attached media asset. Only used when node_id is set.
audio_urlNoAudio track URL. Video duration matches audio duration.
transitionNo
style_presetNoKen Burns style preset: cinematic, dramatic, calm, documentary, dreamy, moody, energetic
random_effectsNoShuffle Ken Burns effects across image items (default: true)
background_colorNoBackground color (default '#000')
duration_secondsNoVideo duration in seconds (required if no audio_url)
background_audio_urlNoAmbience/music bed under the narration (crowd noise, room tone). Loops to fill the video. Find tracks with clipform_search_music.
background_audio_volumeNoBackground bed volume 0-1 (default 0.15 - sits under speech)

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
job_idNoPresent when status is 'rendering' - pass to check_render
statusYes'rendering' when wait:false (poll check_render); 'complete' with a public_url when wait:true
attachedNoTrue when node_id was provided and the render was attached to the node automatically (present once the attach outcome is known).
public_urlNoPresent when status is 'complete' - attach via upload_media_asset then attach_node_media (fit_media: true), unless node_id was set (auto-attached)
republishedNoTrue when the form was live and was republished to include this media.
attach_errorNoPresent when node_id was provided but auto-attach failed - the render itself still succeeded.
media_asset_idNoThe workspace media asset created from this render, when attached.
duration_secondsNo

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. Changed3 schema fields changed
    • changedInput schema / properties / captions / items / properties / words / description
      Previous value: -"Per-word timestamps within the segment"New value: +"Per-word timestamps within the segment. Required - copy the full array from clipform_generate_tts verbatim."
    • addedInput schema / properties / captions / items / properties / words / minItems
      Added value: +1
    • changedInput schema / properties / captions / items / required
      Previous value: -[
      -  "start",
      -  "end",
      -  "text"
      -]New value: +[
      +  "start",
      +  "end",
      +  "text",
      +  "words"
      +]
  2. Changed3 schema fields changed
    • removedInput schema / additionalProperties
      Removed value: -false
    • addedInput schema / properties / context
      Added value: +{
      +  "description": "Describe the user's underlying goal in one sentence - not the tool you're calling.",
      +  "type": "string"
      +}
    • changedInput schema / required
      Previous value: -[
      -  "items"
      -]New value: +[
      +  "items",
      +  "context"
      +]
  3. Changed8 schema fields changed
    • addedInput schema / properties / captions
      Added value: +{
      +  "description": "Word-level captions from clipform_generate_tts - carried onto the attached media asset. Only used when node_id is set.",
      +  "items": {
      +    "additionalProperties": false,
      +    "properties": {
      +      "end": {
      +        "description": "Segment end time in seconds",
      +        "type": "number"
      +      },
      +      "start": {
      +        "description": "Segment start time in seconds",
      +        "type": "number"
      +      },
      +      "text": {
      +        "description": "Full segment text",
      +        "type": "string"
      +      },
      +      "words": {
      +        "description": "Per-word timestamps within the segment",
      +        "items": {
      +          "additionalProperties": false,
      +          "properties": {
      +            "end": {
      +              "type": "number"
      +            },
      +            "start": {
      +              "type": "number"
      +            },
      +            "word": {
      +              "type": "string"
      +            }
      +          },
      +          "required": [
      +            "word",
      +            "start",
      +            "end"
      +          ],
      +          "type": "object"
      +        },
      +        "type": "array"
      +      }
      +    },
      +    "required": [
      +      "start",
      +      "end",
      +      "text"
      +    ],
      +    "type": "object"
      +  },
      +  "type": "array"
      +}
    • addedInput schema / properties / form_id
      Added value: +{
      +  "description": "The form UUID (required when node_id is set).",
      +  "format": "uuid",
      +  "type": "string"
      +}
    • addedInput schema / properties / node_id
      Added value: +{
      +  "description": "Form node to attach this render to automatically once it completes - skips the manual clipform_upload_media_asset + clipform_attach_node_media steps and republishes the form if it's currently live. Requires form_id.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / attach_error
      Added value: +{
      +  "description": "Present when node_id was provided but auto-attach failed - the render itself still succeeded.",
      +  "type": "string"
      +}
    • addedOutput schema / properties / attached
      Added value: +{
      +  "description": "True when node_id was provided and the render was attached to the node automatically (present once the attach outcome is known).",
      +  "type": "boolean"
      +}
    • addedOutput schema / properties / media_asset_id
      Added value: +{
      +  "description": "The workspace media asset created from this render, when attached.",
      +  "type": "string"
      +}
    • changedOutput schema / properties / public_url / description
      Previous value: -"Present when status is 'complete' - attach via upload_media_asset then attach_node_media (fit_media: true)"New value: +"Present when status is 'complete' - attach via upload_media_asset then attach_node_media (fit_media: true), unless node_id was set (auto-attached)"
    • addedOutput schema / properties / republished
      Added value: +{
      +  "description": "True when the form was live and was republished to include this media.",
      +  "type": "boolean"
      +}
  4. Changed1 schema field changed
    • changedOutput schema / properties / public_url / description
      Previous value: -"Present when status is 'complete' - attach via upload_node_media (fit_media: true)"New value: +"Present when status is 'complete' - attach via upload_media_asset then attach_node_media (fit_media: true)"
  5. Changed1 schema field changed
    • changedOutput schema / (root)
      Previous value: -nullNew value: +{
      +  "$schema": "http://json-schema.org/draft-07/schema#",
      +  "additionalProperties": false,
      +  "properties": {
      +    "duration_seconds": {
      +      "type": "number"
      +    },
      +    "job_id": {
      +      "description": "Present when status is 'rendering' - pass to check_render",
      +      "type": "string"
      +    },
      +    "public_url": {
      +      "description": "Present when status is 'complete' - attach via upload_node_media (fit_media: true)",
      +      "type": "string"
      +    },
      +    "status": {
      +      "description": "'rendering' when wait:false (poll check_render); 'complete' with a public_url when wait:true",
      +      "enum": [
      +        "rendering",
      +        "complete"
      +      ],
      +      "type": "string"
      +    }
      +  },
      +  "required": [
      +    "status"
      +  ],
      +  "type": "object"
      +}
  6. First observed

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses key behaviors: fixed 9:16 720x1280 output with Ken Burns effects, muted video clips by default, duration derivation from audio_url, immediate job ID vs blocking URL, attach-on-complete behavior including form republishing, and expected render latency (15-120 seconds per render). These materially inform invocation decisions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but highly structured: purpose, use cases, workflow guidance, sibling-tool decision rules, and node-attachment caveats each occupy purposeful paragraphs. Every section adds operational value rather than restating schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter tool with nested objects and an output schema, this description covers the important operational context: return modes, rendering time, parallel patterns, sibling selection, and side effects. Combined with 94% schema coverage and an output schema, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Although schema coverage is already high (94%), the description adds critical semantics not in the schema: the wait=false parallel workflow, node_id auto-attachment replacing manual upload/attach steps, captions only having effect when node_id is set, and audio-driven duration. This goes well beyond repeating parameter names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Generate a video from images, video clips, or both, synced to an audio track.' It names concrete use cases (narrated question backgrounds, topic visualisations) and distinguishes itself from clipform_render_video_template and clipform_render_composition, so an agent can identify the correct tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use guidance and names alternatives: use a video template for recognisable form beats, use this tool for narrated/audio-synced media montages, and use clipform_render_composition only when neither fits. It also gives a montage disambiguation rule and multi-render workflow advice (wait: false, then poll clipform_check_render).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

A4.1/5.0
Disambiguation4/5

Most tools have clearly distinct purposes: form CRUD, node management, media upload/attach, rendering, search, and guidance retrieval are all separable. The main overlap is among the three render tools (clipform_generate_video, clipform_render_video_template, clipform_render_composition), but their descriptions include explicit disambiguation guidance, so an agent can correctly choose. clipform_get_guide and clipform_get_workflow are also similar but clearly differentiated.

Naming Consistency4/5

Tool names follow a consistent clipform_<verb>_<noun> pattern throughout, e.g., clipform_create_form, clipform_add_node, clipform_update_node, clipform_delete_node. Minor deviations exist: clipform_whoami is not verb_noun, and get_more_tools lacks the clipform_ prefix, but these are edge cases and the overall convention is highly predictable.

Tool Count3/5

34 tools is on the heavy side for a single MCP server. The server covers a broad domain (form creation, node editing, media management, video rendering, TTS, search, guidance, imports, responses), so the count is defensible, but it is above the typical well-scoped range and may add navigation overhead.

Completeness5/5

The tool surface covers the full lifecycle: create/read/update/delete forms and nodes, media upload/attach/delete, multiple render paths with status checking, TTS generation, music/image/video search, form import, response retrieval, and workflow/guide knowledge. The main gap is lack of a direct branching-logic editor (option-based branching is only in the dashboard), but the API consciously documents that limitation and the rest of the lifecycle is complete.