transcriptor-mcp
An MCP server (stdio + HTTP/SSE) that fetches video transcripts/subtitles via yt-dlp, with pagination for large responses. Supports YouTube, Twitter/X, Instagram, TikTok, Twitch, Vimeo, Facebook, Bilibili, VK, Dailymotion. Whisper fallback β transcribes audio when subtitles are unavailable (local or OpenAI API). Works with Cursor and other MCP host
Documentation
π¬ Now your AI assistant can watch videos!
Connect one server. Then ask Claude, ChatGPT or etc about a video: the transcript, the chapters, the metadata, or a single frame. It works with 11 platforms, not only YouTube.
Connect Β· What to ask Β· Widgets Β· Platforms Β· Self-host
β‘ Connect in 30 seconds
The hosted endpoint is:
https://transcriptor.gateway.mcpal.io/mcp
π±οΈ One click
β¨οΈ One command, for Claude Code
claude mcp add --transport http transcriptor https://transcriptor.gateway.mcpal.io/mcp
Then run /mcp and approve the sign-in in the browser. After this, claude mcp list shows β Connected.
π§ No terminal
| Client | What to do |
|---|---|
| Claude (web and desktop) | Open Settings β Customize β Connectors. Select Add β Add custom connector, paste https://transcriptor.gateway.mcpal.io/mcp, then select Add. |
| ChatGPT | Open Settings β Security and login and turn on Developer mode. Then open Plugins, select +, and paste https://transcriptor.gateway.mcpal.io/mcp. |
| Codex | Add the block below to ~/.codex/config.toml, then run codex mcp login transcriptor. The CLI, the IDE extension and the ChatGPT desktop app share this file. |
[mcp_servers.transcriptor]
url = "https://transcriptor.gateway.mcpal.io/mcp"
auth = "oauth"
Note: ChatGPT developer mode is available on the web, for paid plans. Some releases show this control as Settings β Apps & Connectors β Advanced.
Note:
codex mcp addregisters stdio servers only, so a hosted server goes intoconfig.toml. See the Codex MCP docs.
π§© Any other MCP client
If your client is not in the list above, add the server with this configuration:
{
"mcpServers": {
"transcriptor": {
"url": "https://transcriptor.gateway.mcpal.io/mcp"
}
}
}
If you want to run the server yourself, read Self-host. The tools are the same and you need no account.
π§° What you can ask
| Ask for this | Tool |
|---|---|
| "Summarize this video for me" | get_transcript |
| "Give me the subtitles as an SRT file" | get_raw_subtitles |
| "Is there a German track for this video?" | get_available_subtitles |
| "Who published this and how many views?" | get_video_info |
| "Go to the part about pricing" | get_video_chapters |
| "Show me the screen at 4:12" | get_video_frame |
| "Get transcripts for the first 5 videos in this playlist" | get_playlist_transcripts |
| "Find recent videos about X" | search_videos (YouTube) |
Long transcripts come in parts. Each response gives a cursor for the next part, so no text is lost.
Full tool reference (input and structured response)
Each tool that takes a video accepts url. This is a link from a supported platform or a plain YouTube ID. Each tool returns content (text for the chat) and structuredContent (typed JSON for your code).
get_transcript
Clean plain text, without timestamps, HTML, or speaker names. The tool finds the type and the language for you.
Response: videoId, type, lang, text, is_truncated, total_length, start_offset, end_offset. When more text is available, the response also has next_cursor.
get_raw_subtitles
Raw SRT or VTT content, in parts.
Input:
typeβofficialorautolangβ a language coderesponse_limitβ default50000, minimum1000, maximum200000next_cursorβ the cursor of the previous response
Response: the fields of get_transcript, plus format (srt or vtt) and content.
get_available_subtitles
Response: official and auto. Each field is a sorted list of language codes. Use this tool first, then give type and lang to the tools above.
get_video_info
Extended metadata from yt-dlp:
- identity β
videoId,title,description,webpageUrl - author β
uploader,uploaderId,channel,channelId,channelUrl - numbers β
duration,uploadDate,viewCount,likeCount,commentCount - classification β
tags,categories,liveStatus,isLive,wasLive,availability - images β
thumbnailandthumbnails
get_video_chapters
Response: chapters. Each item has startTime, endTime, and title. When the video has no chapters, the list is empty.
get_video_frame
Input:
timecodeβ"MM:SS"or"HH:MM:SS.mmm"secondsβ an alternative totimecode. Give one of the two, not bothformatβjpeg(default) orpngwidthβ default1280, maximum1920, never larger than the sourcequalityβ2to31, for jpeg only
Response: an image block, plus timestampSeconds, timestamp, mimeType, sizeBytes, and width. This tool needs ffmpeg. The Docker image includes it.
get_playlist_transcripts
Input:
urlβ a playlist URL, or a watch URL withlist=type,lang,formatβ the same asget_raw_subtitlesplaylistItemsβ a yt-dlp-Ivalue such as1:5,1,3,7, or-1maxItemsβ the maximum number of videos
Response: results. Each item has videoId and text.
search_videos
Input:
queryβ the search textlimitβ default 10, maximum 50offsetβ the number of results to skipuploadDateFilterβhour,today,week,month, oryearresponse_formatβjson(default) ormarkdown
Response: results. Each item has videoId, title, url, duration, uploader, viewCount, and thumbnail.
πΊ Widgets
Four tools have an interactive interface: get_transcript, get_video_info, get_video_frame, and search_videos. Clients that support MCP Apps and the ChatGPT Apps SDK show this interface in the chat. Other clients get the same data as text and JSON.
|
|
|
|
π Platforms
YouTube Β· Twitter/X Β· Instagram Β· TikTok Β· Twitch Β· Vimeo Β· Facebook Β· Bilibili Β· VK Β· Dailymotion Β· Reddit
Each tool that takes a video accepts a link from these 11 platforms. The tool search_videos works with YouTube only, through yt-dlp ytsearch.
The server does not download video or audio files for you. It returns text, metadata, and single frames.
π³ Self-host
The tools are the same as on the hosted endpoint. You need no account.
Run the server with Docker. The image serves Streamable HTTP on port 4200:
docker run --rm -p 4200:4200 artsamsonov/transcriptor-mcp:latest
Then point your client at http://localhost:4200/mcp.
For stdio, give the image an explicit command:
docker run --rm -i artsamsonov/transcriptor-mcp:latest npm run start:mcp
{
"mcpServers": {
"transcriptor": {
"command": "docker",
"args": ["run", "--rm", "-i", "artsamsonov/transcriptor-mcp:latest", "npm", "run", "start:mcp"]
}
}
}
The server starts with no environment variables. Each variable below is optional.
| Variable | Default | Function |
|---|---|---|
MCP_PORT and MCP_HOST | 4200 and 0.0.0.0 | The HTTP listener |
COOKIES_FILE_PATH | β | A Netscape cookies file for videos that need an account. See cookies.example.txt |
WHISPER_MODE | off | Set local or api to transcribe the audio when a video has no subtitles. Then set WHISPER_BASE_URL or WHISPER_API_KEY |
CACHE_MODE | off | Set redis and CACHE_REDIS_URL to cache subtitles and metadata |
YT_DLP_* | β | Timeouts, proxy, and JS runtimes. See .env.example |
The same port serves GET /health and GET /metrics. The metrics are in Prometheus format and include the mcp_* counters.
Transport, REST API, and development
Transport. The server accepts POST /mcp only. GET and DELETE return 405. The server is stateless and sends no Mcp-Session-Id.
The Node process does not check bearer tokens. Put a reverse proxy or a gateway in front of it for authentication and TLS. The hosted endpoint works this way.
REST API. A second image gives the same extraction over plain HTTP:
docker run --rm -p 3000:3000 artsamsonov/transcriptor-mcp-api:latest
The Swagger interface is at http://localhost:3000/docs. For a full stack with the API and the MCP server, read docker-compose.example.yml.
Development.
npm ci
npm run build
npm run dev:mcp # stdio, hot reload
npm run dev:mcp:http # Streamable HTTP, hot reload
npm test
You need Node.js 20 or later, and yt-dlp in your PATH. Frame capture also needs ffmpeg. Other scripts: lint, type-check, format, test:coverage, test:e2e:api, and test:e2e:mcp.
Releases. The version comes from package.json at runtime, through src/version.ts. Change this version, move the [Unreleased] entries of the changelog into the new version, then push a v* tag. CI builds both images and publishes the MCP Registry entry from server.json.
Layout. src/mcp.ts (stdio entry), src/mcp-http.ts (Streamable HTTP), src/mcp-core.ts (tools, prompts, widgets), src/youtube.ts (yt-dlp), src/whisper.ts, src/cache.ts, src/index.ts (REST API), load/ (k6), and src/e2e/ (Docker smoke tests).
π€ Contributing
Pull requests are welcome. Fork the repository, make a branch, and make sure that npm test and npm run lint pass. Then open a pull request.
βοΈ Legal
The hosted endpoint at transcriptor.gateway.mcpal.io is governed by the Terms of Service and the Privacy Policy.
A server you host yourself is not covered by those documents. It is governed by the MIT License only.
π License
MIT Β© 2026 samson-art. Read LICENSE.
π¬ Support
Issues Β· GitHub profile Β· LinkedIn