transcriptor-mcp

Server MCP (stdio + HTTP/SSE) yang mengambil transkrip/subtitle video melalui yt-dlp, dengan paginasi untuk respons besar. Mendukung YouTube, Twitter/X, Instagram, TikTok, Twitch, Vimeo, Facebook, Bilibili, VK, Dailymotion. Cadangan Whisper โ€” mentranskripsikan audio saat subtitle tidak tersedia (lokal atau API OpenAI). Berfungsi dengan Cursor dan host MCP lainnya.

Dokumentasi

Transcriptor MCP

๐ŸŽฌ Now your AI assistant can watch videos!

Connect one server. Then ask Claude, ChatGPT or etc about a video: the transcript, the chapters, the metadata, or a single frame. It works with 11 platforms, not only YouTube.

MCP Registry Docker MCP Apps License

Connect ยท What to ask ยท Widgets ยท Platforms ยท Self-host


โšก Connect in 30 seconds

The hosted endpoint is:

https://gateway.mcpal.io/mcp/transcriptor

๐Ÿ–ฑ๏ธ One click

Add to Cursor Install in VS Code

โŒจ๏ธ One command, for Claude Code

claude mcp add --transport http transcriptor https://gateway.mcpal.io/mcp/transcriptor

Then run /mcp and approve the sign-in in the browser. After this, claude mcp list shows โœ” Connected.

๐Ÿงญ No terminal

ClientWhat to do
Claude (web and desktop)Open Settings โ†’ Connectors. Select Add custom connector, paste https://gateway.mcpal.io/mcp/transcriptor, then select Add.
ChatGPTOpen Settings โ†’ Security and login and turn on Developer mode. Then open Plugins, select +, and paste https://gateway.mcpal.io/mcp/transcriptor.

Note: ChatGPT developer mode is available on the web, for paid plans. Some releases show this control as Settings โ†’ Apps & Connectors โ†’ Advanced.

๐Ÿงฉ Any other MCP client

If your client is not in the list above, add the server with this configuration:

{
  "mcpServers": {
    "transcriptor": {
      "url": "https://gateway.mcpal.io/mcp/transcriptor"
    }
  }
}

If you want to run the server yourself, read Self-host. The tools are the same and you need no account.


๐Ÿงฐ What you can ask

Ask for thisTool
"Summarize this video for me"get_transcript
"Give me the subtitles as an SRT file"get_raw_subtitles
"Is there a German track for this video?"get_available_subtitles
"Who published this and how many views?"get_video_info
"Go to the part about pricing"get_video_chapters
"Show me the screen at 4:12"get_video_frame
"Get transcripts for the first 5 videos in this playlist"get_playlist_transcripts
"Find recent videos about X"search_videos (YouTube)

Long transcripts come in parts. Each response gives a cursor for the next part, so no text is lost.

Full tool reference (input and structured response)

Each tool that takes a video accepts url. This is a link from a supported platform or a plain YouTube ID. Each tool returns content (text for the chat) and structuredContent (typed JSON for your code).

get_transcript

Clean plain text, without timestamps, HTML, or speaker names. The tool finds the type and the language for you.

Response: videoId, type, lang, text, is_truncated, total_length, start_offset, end_offset. When more text is available, the response also has next_cursor.

get_raw_subtitles

Raw SRT or VTT content, in parts.

Input:

  • type โ€” official or auto
  • lang โ€” a language code
  • response_limit โ€” default 50000, minimum 1000, maximum 200000
  • next_cursor โ€” the cursor of the previous response

Response: the fields of get_transcript, plus format (srt or vtt) and content.

get_available_subtitles

Response: official and auto. Each field is a sorted list of language codes. Use this tool first, then give type and lang to the tools above.

get_video_info

Extended metadata from yt-dlp:

  • identity โ€” videoId, title, description, webpageUrl
  • author โ€” uploader, uploaderId, channel, channelId, channelUrl
  • numbers โ€” duration, uploadDate, viewCount, likeCount, commentCount
  • classification โ€” tags, categories, liveStatus, isLive, wasLive, availability
  • images โ€” thumbnail and thumbnails

get_video_chapters

Response: chapters. Each item has startTime, endTime, and title. When the video has no chapters, the list is empty.

get_video_frame

Input:

  • timecode โ€” "MM:SS" or "HH:MM:SS.mmm"
  • seconds โ€” an alternative to timecode. Give one of the two, not both
  • format โ€” jpeg (default) or png
  • width โ€” default 1280, maximum 1920, never larger than the source
  • quality โ€” 2 to 31, for jpeg only

Response: an image block, plus timestampSeconds, timestamp, mimeType, sizeBytes, and width. This tool needs ffmpeg. The Docker image includes it.

get_playlist_transcripts

Input:

  • url โ€” a playlist URL, or a watch URL with list=
  • type, lang, format โ€” the same as get_raw_subtitles
  • playlistItems โ€” a yt-dlp -I value such as 1:5, 1,3,7, or -1
  • maxItems โ€” the maximum number of videos

Response: results. Each item has videoId and text.

search_videos

Input:

  • query โ€” the search text
  • limit โ€” default 10, maximum 50
  • offset โ€” the number of results to skip
  • uploadDateFilter โ€” hour, today, week, month, or year
  • response_format โ€” json (default) or markdown

Response: results. Each item has videoId, title, url, duration, uploader, viewCount, and thumbnail.


๐Ÿ“บ Widgets

Four tools have an interactive interface: get_transcript, get_video_info, get_video_frame, and search_videos. Clients that support MCP Apps and the ChatGPT Apps SDK show this interface in the chat. Other clients get the same data as text and JSON.

The search_videos widget: a carousel of result cards with thumbnails, durations, and view counts

search_videos ยท "model context protocol MCP server production"

The get_video_frame widget: one captured frame with step controls and a timecode field

get_video_frame ยท an architecture slide at 3:30

The get_transcript widget: a video card above a searchable list of timed captions

get_transcript ยท a 3-minute MCP explainer, official captions

The get_video_info widget: thumbnail, channel, views, likes, description, and a subtitle language picker

get_video_info ยท channel, views, likes, and 169 caption languages


๐ŸŒ Platforms

YouTube ยท Twitter/X ยท Instagram ยท TikTok ยท Twitch ยท Vimeo ยท Facebook ยท Bilibili ยท VK ยท Dailymotion ยท Reddit

Each tool that takes a video accepts a link from these 11 platforms. The tool search_videos works with YouTube only, through yt-dlp ytsearch.

The server does not download video or audio files for you. It returns text, metadata, and single frames.


๐Ÿณ Self-host

The tools are the same as on the hosted endpoint. You need no account.

Run the server with Docker. The image serves Streamable HTTP on port 4200:

docker run --rm -p 4200:4200 artsamsonov/transcriptor-mcp:latest

Then point your client at http://localhost:4200/mcp.

For stdio, give the image an explicit command:

docker run --rm -i artsamsonov/transcriptor-mcp:latest npm run start:mcp
{
  "mcpServers": {
    "transcriptor": {
      "command": "docker",
      "args": ["run", "--rm", "-i", "artsamsonov/transcriptor-mcp:latest", "npm", "run", "start:mcp"]
    }
  }
}

The server starts with no environment variables. Each variable below is optional.

VariableDefaultFunction
MCP_PORT and MCP_HOST4200 and 0.0.0.0The HTTP listener
COOKIES_FILE_PATHโ€”A Netscape cookies file for videos that need an account. See cookies.example.txt
WHISPER_MODEoffSet local or api to transcribe the audio when a video has no subtitles. Then set WHISPER_BASE_URL or WHISPER_API_KEY
CACHE_MODEoffSet redis and CACHE_REDIS_URL to cache subtitles and metadata
YT_DLP_*โ€”Timeouts, proxy, and JS runtimes. See .env.example

The same port serves GET /health and GET /metrics. The metrics are in Prometheus format and include the mcp_* counters.

Transport, REST API, and development

Transport. The server accepts POST /mcp only. GET and DELETE return 405. The server is stateless and sends no Mcp-Session-Id.

The Node process does not check bearer tokens. Put a reverse proxy or a gateway in front of it for authentication and TLS. The hosted endpoint works this way.

REST API. A second image gives the same extraction over plain HTTP:

docker run --rm -p 3000:3000 artsamsonov/transcriptor-mcp-api:latest

The Swagger interface is at http://localhost:3000/docs. For a full stack with the API and the MCP server, read docker-compose.example.yml.

Development.

npm ci
npm run build
npm run dev:mcp        # stdio, hot reload
npm run dev:mcp:http   # Streamable HTTP, hot reload
npm test

You need Node.js 20 or later, and yt-dlp in your PATH. Frame capture also needs ffmpeg. Other scripts: lint, type-check, format, test:coverage, test:e2e:api, and test:e2e:mcp.

Releases. The version comes from package.json at runtime, through src/version.ts. Change this version, move the [Unreleased] entries of the changelog into the new version, then push a v* tag. CI builds both images and publishes the MCP Registry entry from server.json.

Layout. src/mcp.ts (stdio entry), src/mcp-http.ts (Streamable HTTP), src/mcp-core.ts (tools, prompts, widgets), src/youtube.ts (yt-dlp), src/whisper.ts, src/cache.ts, src/index.ts (REST API), load/ (k6), and src/e2e/ (Docker smoke tests).


๐Ÿค Contributing

Pull requests are welcome. Fork the repository, make a branch, and make sure that npm test and npm run lint pass. Then open a pull request.

โš–๏ธ Legal

The hosted endpoint at gateway.mcpal.io is governed by the Terms of Service and the Privacy Policy.

A server you host yourself is not covered by those documents. It is governed by the MIT License only.

๐Ÿ“„ License

MIT ยฉ 2026 samson-art. Read LICENSE.

๐Ÿ’ฌ Support

Issues ยท GitHub profile ยท LinkedIn