VideoGen MCP
通过官方VideoGen API,从AI代理创建视频、图像、配音、音乐和虚拟形象。
托管 MCP 服务器
npx add-mcp 'https://mcp.videogen.io/mcp'可安装到 Claude Code、Codex、Cursor 等客户端
文档
For clean Markdown of any page, append .md to the page URL. For a complete documentation index, see https://docs.videogen.io/llms.txt. For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.videogen.io/_mcp/server.
MCP server
Connect the VideoGen MCP server (hosted remote or local) that exposes every VideoGen API capability as a tool for Cursor, Claude Desktop, Windsurf, and other MCP clients.
The VideoGen MCP server is a Model Context Protocol server that exposes the full VideoGen API to any MCP client. Point your agent at it and it can generate videos from scripts, produce images, voiceovers, music, and avatars, upload files, export and remix projects, and manage runs, all authenticated with your own API key. It comes in two transports that expose the same tools: a hosted remote server (recommended) and a local server.
Note
This is the API MCP server, which executes real API calls. It is distinct from the hosted documentation MCP at
https://docs.videogen.io/_mcp/server, which is read-only and only answers questions about the docs. See Documentation MCP below.
Connect
Both transports wrap the official @videogen/sdk and expose the same tools. Get an API key from app.videogen.io/api first: every tool call runs as the team that owns the key.
Remote (recommended)
The hosted server needs nothing to install or update. Point your MCP client at the endpoint and send your key as a bearer token:
{
"mcpServers": {
"videogen": {
"url": "https://mcp.videogen.io/mcp",
"headers": {
"Authorization": "Bearer sk_videogen_live_..."
}
}
}
}
The hosted server is stateless and multi-tenant. Your key is read from the request header, forwarded only to the VideoGen API, and never stored. Because it runs in the cloud and has no access to your machine's filesystem, upload_file is not available on the remote server. Upload files directly through the VideoGen API (request a presigned upload URL and PUT the bytes yourself), then pass the returned vg_file_... id to other tools — or use the local server for files on your machine.
One-click install
Install the hosted server in one click, then authenticate with Sign in with VideoGen (or add your API key as an Authorization header afterward):
- Cursor: Add VideoGen to Cursor
- VS Code: Add VideoGen to VS Code
You can also install it from the command line:
code --add-mcp '{"name":"videogen","type":"http","url":"https://mcp.videogen.io/mcp"}'
Local (stdio)
The local server runs as a subprocess that your MCP client launches with npx, so there is nothing to install ahead of time. It reads your key from the VIDEOGEN_API_KEY environment variable:
{
"mcpServers": {
"videogen": {
"command": "npx",
"args": ["-y", "@videogen/mcp"],
"env": {
"VIDEOGEN_API_KEY": "sk_videogen_live_..."
}
}
}
}
Your key stays on your machine. It is passed directly to the local server process and never sent anywhere except the VideoGen API.
| Variable | Required | Default | Description |
|---|---|---|---|
VIDEOGEN_API_KEY | local only | — | Your VideoGen API key (local server). On the remote server the key travels in the Authorization: Bearer header instead. |
VIDEOGEN_BASE_URL | no | https://api.videogen.io | Override the upstream API base URL (e.g. for local development). |
Other MCP clients
Windsurf, Cline, Goose, Claude Desktop, and most other MCP clients accept the same mcpServers JSON shown above (remote or local). Add the block to that client's MCP config, then reload:
| Client | Where to add it |
|---|---|
| Cursor | Settings → MCP → Add, or the one-click link above |
| VS Code | .vscode/mcp.json, or the one-click link above |
| Claude Desktop | Settings → Developer → Edit Config (claude_desktop_config.json) |
| Claude (web/mobile) | Settings → Connectors → Add custom connector — see Connect to Claude |
| Windsurf | ~/.codeium/windsurf/mcp_config.json |
| Cline | MCP Servers panel → Configure MCP Servers |
| Goose | Settings → Extensions → Add, or ~/.config/goose/config.yaml |
Note
Clients that support remote MCP over OAuth (Cursor, VS Code, Claude) can connect with just the server URL and authenticate via Sign in with VideoGen. Clients without remote/OAuth support use the local (stdio) block with a
VIDEOGEN_API_KEY.
Sign in with VideoGen (OAuth)
MCP hosts that support OAuth — including Claude — can connect to the hosted server without an API key. Each user signs in through Sign in with VideoGen and the host calls the server on their behalf. The server advertises its authorization server via protected-resource metadata and supports Dynamic Client Registration (RFC 7591), so these hosts discover the auth server and register themselves automatically — there is nothing to paste but the server URL.
Use https://mcp.videogen.io/mcp for Cursor, Claude, VS Code, and other hosts that follow transport-level OAuth challenges.
Note
Cursor: Prefer the remote config with an
Authorization: BearerAPI key above. Cursor's OAuth client strips path components from authorization-server URLs (it does not follow RFC 8414 for path issuers), which breaks Sign in with VideoGen against Supabase Auth (…/auth/v1). The hosted/mcpendpoint works around this by advertising a pathless authorization server onmcp.videogen.io; if OAuth still fails after a Cursor update, use an API key.
Connect to Claude
- In Claude (claude.ai, Claude Desktop, or the mobile apps), open Settings → Connectors and click Add custom connector.
- Enter the MCP server URL
https://mcp.videogen.io/mcp. Leave the OAuth Client ID / Secret fields in Advanced settings blank — VideoGen supports Dynamic Client Registration, so Claude registers itself automatically. - Click Connect and approve the Sign in with VideoGen consent.
Each host uses its own OAuth redirect (callback) URL. VideoGen's authorization server accepts these automatically through Dynamic Client Registration, so you don't need to configure them — they're listed here for reference:
| Host | OAuth callback (redirect) URL |
|---|---|
| Claude — web, Desktop, mobile, Cowork | https://claude.ai/api/mcp/auth_callback |
| Claude Code | http://localhost/callback and http://127.0.0.1/callback (loopback, port-agnostic) |
Signed-in users can review and revoke connected apps at any time from the developer dashboard.
Long-running operations
Workflows, media tools, and project exports are asynchronous. The MCP server wraps each of these as a composite tool that starts the operation and, by default, polls until it reaches a terminal state (succeeded, failed, or cancelled) before returning the finished result. Every composite tool accepts these optional control fields:
| Field | Type | Default | Description |
|---|---|---|---|
wait | boolean | true | Block until the operation reaches a terminal state. Set false to return immediately with the run/execution id. |
pollIntervalMs | number | — | How often to poll while waiting, in milliseconds. |
timeoutMs | number | — | Maximum time to wait for a terminal state before giving up, in milliseconds. |
When wait is false, use the matching get_* tool (e.g. get_workflow_run, get_tool_execution) to poll the returned id yourself.
Guidance resources
The API MCP also exposes operational guidance as static MCP resources (guidance://getting-started, guidance://async-tasks, guidance://workflows, guidance://tools-vs-workflows) plus matching no-arg tools (get_getting_started_guidance, get_async_tasks_guidance, get_workflows_guidance, get_tools_vs_workflows_guidance). Agents should call those tools when choosing workflows vs media tools, handling async polling on the hosted server, or following the run → remix → export flow. This is separate from the documentation MCP, which serves the full Fern docs site read-only.
Tool reference
The server registers 49 tools: 45 API tools plus four guidance tools. The remote transport also provides a hosted-only open_uploader widget. Ids are vg-prefixed opaque strings (vg_file_..., vg_enti_..., vg_work_..., vg_tool_..., vg_voic_...).
Workflows (end-to-end video)
| Tool | Description | Key parameters |
|---|---|---|
script_to_video | Turn a script into a finished narrated video with visuals and captions. The script is narrated verbatim. Composite (waits by default). | script, visualStyle, aspectRatio, visualPacing, quality, language, voiceId, voiceSpeed, actorEntityId, avatarQuality, featuredBRollFileIds, workflowAgentContext, scenes, remixActions, + poll controls |
voiceover_to_video | Build a narrated video from an already-uploaded voiceover audio file. Upload with upload_file first. Composite. | fileId, visualStyle, aspectRatio, visualPacing, quality, language, captionStyle, logoFileId, workflowAgentContext, scenes, remixActions, + poll controls |
slideshow_to_video | Build a narrated video from an already-uploaded PDF or slideshow file. Upload with upload_file first. Composite. | fileId, slideScripts, aspectRatio, language, voiceId, voiceSpeed, actorEntityId, avatarQuality, captionStyle, logoFileId, remixActions, + poll controls |
storyboard_to_video | Build a video from a structured storyboard of scenes. Composite. | scenes, actorEntityIds, productEntityIds, defaultGeneration, defaultDurationSeconds, quality, aspectRatio, workflowAgentContext, remixActions, + poll controls |
prompt_to_video_clip | Create a project and generate one short AI video clip from a text prompt (opening frame, then animate). Composite. Does not accept remixActions. | prompt, imageFileIds, durationSeconds, aspectRatio, quality, + poll controls |
list_workflow_runs | List workflow runs, most recent first. | cursor, limit, selfOnly |
get_workflow_run | Fetch the current status and result of a single workflow run. | workflowRunId |
cancel_workflow_run | Request cancellation of an in-progress workflow run. | workflowRunId |
Provide at least two remixActions (e.g. ENABLE_CAPTIONS + SET_BACKGROUND_MUSIC) for a polished result. Available remix action types: SET_BACKGROUND_MUSIC, SET_LOGO, ENABLE_CAPTIONS, DISABLE_CAPTIONS, ADD_TRANSITIONS, ADD_ZOOM, RESIZE_PROJECT, CLEAN_UP_TRANSCRIPT, CONVERT_IMAGES_TO_VIDEOS.
Media tools
| Tool | Description | Key parameters |
|---|---|---|
generate_image | Generate an image from a text prompt, optionally conditioned on source images. Composite. | prompt, quality, imageFileIds, aspectRatio, watermarkMode, numResults, isOutputTemporary, + poll controls |
generate_video_clip | Generate a video clip from a text prompt, source images, or source videos. Composite. | quality (LOW | STANDARD | HIGH | MAX), prompt, startFrameFileId, imageFileIds, videoFileIds, audioFileIds, spokenDialogue, voiceDescription, generateAudio, suppressBackgroundMusic, durationSeconds, aspectRatio, watermarkMode, numResults, isOutputTemporary, + poll controls |
text_to_speech | Convert text into spoken audio using a selectable voice. Composite. | ttsText, voiceId, speechLanguageCode, pronunciationReplacements, autoExpandPronunciationReplacements, voiceSpeed, numResults, isOutputTemporary, + poll controls |
generate_sound_effect | Generate a sound effect from a text prompt. Composite. | prompt, durationSeconds, promptInfluence, numResults, isOutputTemporary, + poll controls |
generate_music | Generate a music track from a text prompt. Composite. | prompt, numResults, isOutputTemporary, + poll controls |
generate_motion_graphic | Generate an animated motion graphic video from a text prompt (experimental, agentic). Outputs a transparent WebM overlay by default; set transparentBackground to false for an opaque MP4. Composite. | prompt, fileIds, durationSeconds, aspectRatio, transparentBackground |
generate_avatar | Generate a talking-head avatar video from an actor entity and an uploaded audio file. Composite. | actorEntityId, audioFileId, avatarQuality, watermarkMode, numResults, isOutputTemporary, + poll controls |
vectorize_image | Convert a raster image into a vector (SVG). Composite. | imageFileId, watermarkMode, numResults, isOutputTemporary, + poll controls |
remove_image_background | Remove the background from an image. Composite. | imageFileId, watermarkMode, numResults, isOutputTemporary, + poll controls |
remove_video_background | Remove the background from a video. Composite. | videoFileId, watermarkMode, numResults, isOutputTemporary, + poll controls |
upscale_image | Increase the resolution of an image. Composite. | imageFileId, watermarkMode, numResults, isOutputTemporary, + poll controls |
upscale_video | Increase the resolution of a video. Composite. | videoFileId, watermarkMode, numResults, isOutputTemporary, + poll controls |
image_3d_effect | Add 3D parallax motion to a still image, producing a video. Composite. | imageFileId, watermarkMode, numResults, isOutputTemporary, + poll controls |
list_tool_executions | List past tool executions, most recent first. | cursor, limit, selfOnly |
get_tool_execution | Fetch the current status and results of a single tool execution. | toolExecutionId |
cancel_tool_execution | Request cancellation of an in-progress tool execution. | toolExecutionId |
Projects
| Tool | Description | Key parameters |
|---|---|---|
list_projects | List projects. API-created only by default; pass includeUiProjects for dashboard projects too. | cursor, limit, selfOnly, includeUiProjects |
get_project | Fetch metadata and the shareable URL for a single project. | projectId |
export_project | Export a project to an MP4. Composite — waits until the download URL is ready by default. | projectId, quality (STANDARD | HIGH | FULL_HIGH | ULTRA_HIGH), + poll controls |
get_project_export | Fetch the current status of a project export. Poll until status is succeeded, failed, or cancelled. | projectId, exportId |
remix_project | Apply remix actions (music, logo, captions, transitions, natural-language edits) to an existing project. | projectId, remixActions, saveAsNewProject |
list_project_remix_actions | List the status of remix actions applied to a project. | projectId |
Files
| Tool | Description | Key parameters |
|---|---|---|
upload_file | Upload a local file and wait until it is processed. Returns the file id (vg_file_...). Use it for voiceover_to_video, slideshow_to_video, logos, or B-roll. To upload a remote asset, download it first and pass its local path. Local (stdio) server only. | filePath, displayName, type (IMAGE | VIDEO | AUDIO) |
create_file_upload | Create a pending file and return a presigned upload URL (same as POST /v1/files/upload). PUT the bytes yourself, then poll with get_file. | displayName, type, isTemporary, transcript |
get_file | Fetch a file by id with freshly hydrated signed URLs for thumbnail, preview, and download renditions. | fileId |
list_files | List files visible to the current API key. | cursor, limit |
Hosted remote only:
| Tool | Description | Key parameters |
|---|---|---|
open_uploader | Open the hosted upload widget so a human can pick files in the browser (ChatGPT MCP Apps). Not counted in the 46 API tools above. | — |
Entities
| Tool | Description | Key parameters |
|---|---|---|
list_entities | List built-in catalog entities plus team ACTOR, PRODUCT, VISUAL_STYLE, and SLIDESHOW_THEME entities. Built-in rows have isBuiltIn: true and cannot be updated or archived. | entityType, cursor, limit |
create_entity | Create an entity. Attach at least one image afterward with add_entity_reference. | entityType (ACTOR | PRODUCT | VISUAL_STYLE | SLIDESHOW_THEME), name, description |
get_entity | Fetch one entity and its reference images. | entityId |
update_entity | Update an entity's display name and/or description. | entityId, name, description |
archive_entity | Archive an entity so it no longer appears in lists or pickers. | entityId |
add_entity_reference | Attach an uploaded image (vg_file_...) as a reference. Use isDefault: true for the primary thumbnail. | entityId, fileId, description, isDefault |
remove_entity_reference | Detach a reference image from an entity. | entityId, fileId |
Guidance
| Tool | Description | Key parameters |
|---|---|---|
get_getting_started_guidance | Auth, verify with get_me, id conventions, MCP vs SDK. Mirrors guidance://getting-started. | — |
get_async_tasks_guidance | Polling, hosted wait caps, webhooks outside MCP. Mirrors guidance://async-tasks. | — |
get_workflows_guidance | Run → remix → export and each workflow tool. Mirrors guidance://workflows. | — |
get_tools_vs_workflows_guidance | Single-asset tools (automatic model routing) vs full video workflows. Mirrors guidance://tools-vs-workflows. | — |
Resources & account
| Tool | Description | Key parameters |
|---|---|---|
list_tts_voices | List available text-to-speech voices for narration, text_to_speech, and workflows. | cursor, limit, includeDeprecatedVoices, query |
list_languages | List supported languages for narration and captions. | query |
get_me | Fetch the account and team behind the API key (apiKeyId, apiKeyNickname, email, displayName, teamId). Use it as a connection test. | — |
get_app_deep_link | Build a VideoGen app URL that opens a modal or navigates after sign-in (upgrade, buy credits, invite teammates, feedback, integrations, account settings, or a NAVIGATE destination). Prefer this when the user must complete something in the VideoGen UI that MCP tools cannot do inline. No API key required. | action, plus action-specific fields (destination, articleSlug, provider) |
MCP tool to REST endpoint
Each MCP tool maps to one or more VideoGen REST endpoints. Composite tools call the start endpoint and then poll the get endpoint until terminal.
| MCP tool | REST endpoint(s) |
|---|---|
script_to_video | POST /v1/workflows/script-to-video, then GET /v1/workflows/runs/{workflowRunId} |
voiceover_to_video | POST /v1/workflows/voiceover-to-video, then GET /v1/workflows/runs/{workflowRunId} |
slideshow_to_video | POST /v1/workflows/slideshow-to-video, then GET /v1/workflows/runs/{workflowRunId} |
storyboard_to_video | POST /v1/workflows/storyboard-to-video, then GET /v1/workflows/runs/{workflowRunId} |
prompt_to_video_clip | POST /v1/workflows/prompt-to-video-clip, then GET /v1/workflows/runs/{workflowRunId} |
list_workflow_runs | GET /v1/workflows/runs |
get_workflow_run | GET /v1/workflows/runs/{workflowRunId} |
cancel_workflow_run | POST /v1/workflows/runs/{workflowRunId}/cancel |
generate_image | POST /v1/tools/generate-image, then GET /v1/tools/executions/{toolExecutionId} |
generate_video_clip | POST /v1/tools/generate-video-clip, then GET /v1/tools/executions/{toolExecutionId} |
text_to_speech | POST /v1/tools/text-to-speech, then GET /v1/tools/executions/{toolExecutionId} |
generate_sound_effect | POST /v1/tools/generate-sound-effect, then GET /v1/tools/executions/{toolExecutionId} |
generate_music | POST /v1/tools/generate-music, then GET /v1/tools/executions/{toolExecutionId} |
generate_motion_graphic | POST /v1/tools/generate-motion-graphic, then GET /v1/tools/executions/{toolExecutionId} |
generate_avatar | POST /v1/tools/generate-avatar, then GET /v1/tools/executions/{toolExecutionId} |
vectorize_image | POST /v1/tools/vectorize-image, then GET /v1/tools/executions/{toolExecutionId} |
remove_image_background | POST /v1/tools/remove-image-background, then GET /v1/tools/executions/{toolExecutionId} |
remove_video_background | POST /v1/tools/remove-video-background, then GET /v1/tools/executions/{toolExecutionId} |
upscale_image | POST /v1/tools/upscale-image, then GET /v1/tools/executions/{toolExecutionId} |
upscale_video | POST /v1/tools/upscale-video, then GET /v1/tools/executions/{toolExecutionId} |
image_3d_effect | POST /v1/tools/image-3d-effect, then GET /v1/tools/executions/{toolExecutionId} |
list_tool_executions | GET /v1/tools/executions |
get_tool_execution | GET /v1/tools/executions/{toolExecutionId} |
cancel_tool_execution | POST /v1/tools/executions/{toolExecutionId}/cancel |
list_projects | GET /v1/projects |
get_project | GET /v1/projects/{projectId} |
export_project | POST /v1/projects/{projectId}/export, then GET /v1/projects/{projectId}/exports/{exportId} |
get_project_export | GET /v1/projects/{projectId}/exports/{exportId} |
remix_project | POST /v1/projects/{projectId}/remix |
list_project_remix_actions | GET /v1/projects/{projectId}/remix-actions |
upload_file | POST /v1/files/upload, PUT bytes, then GET /v1/files/{fileId} |
create_file_upload | POST /v1/files/upload |
get_file | POST /v1/files/{fileId}/hydrate |
list_files | GET /v1/files |
open_uploader | Hosted upload widget (no direct REST equivalent) |
list_entities | GET /v1/entities |
create_entity | POST /v1/entities |
get_entity | GET /v1/entities/{entityId} |
update_entity | POST /v1/entities/{entityId}/update |
archive_entity | POST /v1/entities/{entityId}/archive |
add_entity_reference | POST /v1/entities/{entityId}/references |
remove_entity_reference | POST /v1/entities/{entityId}/references/remove |
list_tts_voices | GET /v1/resources/tts-voices |
list_languages | GET /v1/resources/languages |
get_me | GET /v1/me |
get_app_deep_link | App deep link (no direct REST equivalent) |
Example prompts
Once the server is connected, ask your agent in natural language. It selects the matching tool automatically:
- "Generate a narrated video from this script with captions and background music" →
script_to_video - "Upload this voiceover and turn it into a video" →
upload_filethenvoiceover_to_video - "Make an image of a red sports car at sunset" →
generate_image - "Create a product entity from this pack-shot image" →
upload_file/open_uploader, thencreate_entity+add_entity_reference - "Read this text aloud in a calm voice" →
list_tts_voicesthentext_to_speech - "Upscale this video to a higher resolution" →
upload_filethenupscale_video - "Export my project as an MP4" →
export_project - "Add background music and a logo to my project" →
remix_project - "Which account and team does this API key belong to?" →
get_me
Documentation MCP
VideoGen also hosts a read-only documentation MCP server that lets AI clients query the API docs in real time. It answers questions about the API but does not call it. Connect it alongside (or instead of) the API server:
{
"mcpServers": {
"videogen-docs": {
"url": "https://docs.videogen.io/_mcp/server"
}
}
}