transcriptor-mcp

Stdio ve HTTP/SSE üzerinden çalışan, yt-dlp ile video transkriptlerini/altyazılarını getiren, büyük yanıtlar için sayfalama desteği sunan bir MCP sunucusu. YouTube, Twitter/X, Instagram, TikTok, Twitch, Vimeo, Facebook, Bilibili, VK, Dailymotion'ı destekler. Whisper yedekleme — altyazı bulunmadığında sesi transkribe eder (yerel veya OpenAI API). Cursor ve diğer MCP ana bilgisayarlarla çalışır.

Dokümantasyon

Transcriptor MCP

🎬 Now your AI assistant can watch videos!

Connect one server. Then ask Claude, ChatGPT or etc about a video: the transcript, the chapters, the metadata, or a single frame. It works with 11 platforms, not only YouTube.

MCP Registry Docker MCP Apps License

Connect · What to ask · Widgets · Platforms · Self-host


⚡ Connect in 30 seconds

The hosted endpoint is:

https://gateway.mcpal.io/mcp/transcriptor

🖱️ One click

Add to Cursor Install in VS Code

⌨️ One command, for Claude Code

claude mcp add --transport http transcriptor https://gateway.mcpal.io/mcp/transcriptor

Then run /mcp and approve the sign-in in the browser. After this, claude mcp list shows ✔ Connected.

🧭 No terminal

ClientWhat to do
Claude (web and desktop)Open Settings → Connectors. Select Add custom connector, paste https://gateway.mcpal.io/mcp/transcriptor, then select Add.
ChatGPTOpen Settings → Security and login and turn on Developer mode. Then open Plugins, select +, and paste https://gateway.mcpal.io/mcp/transcriptor.

Note: ChatGPT developer mode is available on the web, for paid plans. Some releases show this control as Settings → Apps & Connectors → Advanced.

🧩 Any other MCP client

If your client is not in the list above, add the server with this configuration:

{
  "mcpServers": {
    "transcriptor": {
      "url": "https://gateway.mcpal.io/mcp/transcriptor"
    }
  }
}

If you want to run the server yourself, read Self-host. The tools are the same and you need no account.


🧰 What you can ask

Ask for thisTool
"Summarize this video for me"get_transcript
"Give me the subtitles as an SRT file"get_raw_subtitles
"Is there a German track for this video?"get_available_subtitles
"Who published this and how many views?"get_video_info
"Go to the part about pricing"get_video_chapters
"Show me the screen at 4:12"get_video_frame
"Get transcripts for the first 5 videos in this playlist"get_playlist_transcripts
"Find recent videos about X"search_videos (YouTube)

Long transcripts come in parts. Each response gives a cursor for the next part, so no text is lost.

Full tool reference (input and structured response)

Each tool that takes a video accepts url. This is a link from a supported platform or a plain YouTube ID. Each tool returns content (text for the chat) and structuredContent (typed JSON for your code).

get_transcript

Clean plain text, without timestamps, HTML, or speaker names. The tool finds the type and the language for you.

Response: videoId, type, lang, text, is_truncated, total_length, start_offset, end_offset. When more text is available, the response also has next_cursor.

get_raw_subtitles

Raw SRT or VTT content, in parts.

Input:

  • typeofficial or auto
  • lang — a language code
  • response_limit — default 50000, minimum 1000, maximum 200000
  • next_cursor — the cursor of the previous response

Response: the fields of get_transcript, plus format (srt or vtt) and content.

get_available_subtitles

Response: official and auto. Each field is a sorted list of language codes. Use this tool first, then give type and lang to the tools above.

get_video_info

Extended metadata from yt-dlp:

  • identity — videoId, title, description, webpageUrl
  • author — uploader, uploaderId, channel, channelId, channelUrl
  • numbers — duration, uploadDate, viewCount, likeCount, commentCount
  • classification — tags, categories, liveStatus, isLive, wasLive, availability
  • images — thumbnail and thumbnails

get_video_chapters

Response: chapters. Each item has startTime, endTime, and title. When the video has no chapters, the list is empty.

get_video_frame

Input:

  • timecode"MM:SS" or "HH:MM:SS.mmm"
  • seconds — an alternative to timecode. Give one of the two, not both
  • formatjpeg (default) or png
  • width — default 1280, maximum 1920, never larger than the source
  • quality2 to 31, for jpeg only

Response: an image block, plus timestampSeconds, timestamp, mimeType, sizeBytes, and width. This tool needs ffmpeg. The Docker image includes it.

get_playlist_transcripts

Input:

  • url — a playlist URL, or a watch URL with list=
  • type, lang, format — the same as get_raw_subtitles
  • playlistItems — a yt-dlp -I value such as 1:5, 1,3,7, or -1
  • maxItems — the maximum number of videos

Response: results. Each item has videoId and text.

search_videos

Input:

  • query — the search text
  • limit — default 10, maximum 50
  • offset — the number of results to skip
  • uploadDateFilterhour, today, week, month, or year
  • response_formatjson (default) or markdown

Response: results. Each item has videoId, title, url, duration, uploader, viewCount, and thumbnail.


📺 Widgets

Four tools have an interactive interface: get_transcript, get_video_info, get_video_frame, and search_videos. Clients that support MCP Apps and the ChatGPT Apps SDK show this interface in the chat. Other clients get the same data as text and JSON.

The search_videos widget: a carousel of result cards with thumbnails, durations, and view counts

search_videos · "model context protocol MCP server production"

The get_video_frame widget: one captured frame with step controls and a timecode field

get_video_frame · an architecture slide at 3:30

The get_transcript widget: a video card above a searchable list of timed captions

get_transcript · a 3-minute MCP explainer, official captions

The get_video_info widget: thumbnail, channel, views, likes, description, and a subtitle language picker

get_video_info · channel, views, likes, and 169 caption languages


🌍 Platforms

YouTube · Twitter/X · Instagram · TikTok · Twitch · Vimeo · Facebook · Bilibili · VK · Dailymotion · Reddit

Each tool that takes a video accepts a link from these 11 platforms. The tool search_videos works with YouTube only, through yt-dlp ytsearch.

The server does not download video or audio files for you. It returns text, metadata, and single frames.


🐳 Self-host

The tools are the same as on the hosted endpoint. You need no account.

Run the server with Docker. The image serves Streamable HTTP on port 4200:

docker run --rm -p 4200:4200 artsamsonov/transcriptor-mcp:latest

Then point your client at http://localhost:4200/mcp.

For stdio, give the image an explicit command:

docker run --rm -i artsamsonov/transcriptor-mcp:latest npm run start:mcp
{
  "mcpServers": {
    "transcriptor": {
      "command": "docker",
      "args": ["run", "--rm", "-i", "artsamsonov/transcriptor-mcp:latest", "npm", "run", "start:mcp"]
    }
  }
}

The server starts with no environment variables. Each variable below is optional.

VariableDefaultFunction
MCP_PORT and MCP_HOST4200 and 0.0.0.0The HTTP listener
COOKIES_FILE_PATHA Netscape cookies file for videos that need an account. See cookies.example.txt
WHISPER_MODEoffSet local or api to transcribe the audio when a video has no subtitles. Then set WHISPER_BASE_URL or WHISPER_API_KEY
CACHE_MODEoffSet redis and CACHE_REDIS_URL to cache subtitles and metadata
YT_DLP_*Timeouts, proxy, and JS runtimes. See .env.example

The same port serves GET /health and GET /metrics. The metrics are in Prometheus format and include the mcp_* counters.

Transport, REST API, and development

Transport. The server accepts POST /mcp only. GET and DELETE return 405. The server is stateless and sends no Mcp-Session-Id.

The Node process does not check bearer tokens. Put a reverse proxy or a gateway in front of it for authentication and TLS. The hosted endpoint works this way.

REST API. A second image gives the same extraction over plain HTTP:

docker run --rm -p 3000:3000 artsamsonov/transcriptor-mcp-api:latest

The Swagger interface is at http://localhost:3000/docs. For a full stack with the API and the MCP server, read docker-compose.example.yml.

Development.

npm ci
npm run build
npm run dev:mcp        # stdio, hot reload
npm run dev:mcp:http   # Streamable HTTP, hot reload
npm test

You need Node.js 20 or later, and yt-dlp in your PATH. Frame capture also needs ffmpeg. Other scripts: lint, type-check, format, test:coverage, test:e2e:api, and test:e2e:mcp.

Releases. The version comes from package.json at runtime, through src/version.ts. Change this version, move the [Unreleased] entries of the changelog into the new version, then push a v* tag. CI builds both images and publishes the MCP Registry entry from server.json.

Layout. src/mcp.ts (stdio entry), src/mcp-http.ts (Streamable HTTP), src/mcp-core.ts (tools, prompts, widgets), src/youtube.ts (yt-dlp), src/whisper.ts, src/cache.ts, src/index.ts (REST API), load/ (k6), and src/e2e/ (Docker smoke tests).


🤝 Contributing

Pull requests are welcome. Fork the repository, make a branch, and make sure that npm test and npm run lint pass. Then open a pull request.

⚖️ Legal

The hosted endpoint at gateway.mcpal.io is governed by the Terms of Service and the Privacy Policy.

A server you host yourself is not covered by those documents. It is governed by the MIT License only.

📄 License

MIT © 2026 samson-art. Read LICENSE.

💬 Support

Issues · GitHub profile · LinkedIn