tts-generation

作者: google-gemini

使用Interactions API将广播节目脚本转换为语音音频,并对记者声音施加电话效果。

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill tts-generation

TTS Generation

Convert the radio show script into speech audio using the Gemini TTS model via the Interactions API. Apply a telephone bandpass filter to correspondent voices so they sound like phone call-ins, while keeping the host's voice clean studio-quality.

Embedded Script

python3 skills/tts-generation/scripts/generate_tts.py --workspace ./workspace

Arguments

ArgumentDefaultDescription
--workspaceworkspaceRoot workspace directory
--workers8Max parallel TTS worker threads

What it does

  1. Reads the script from {workspace}/data/script.md.
  2. Parses it into individual (speaker, text) turns and assigns voices.
  3. Generates TTS for all turns in parallel using a thread pool (default 8 workers).
  4. Retries each failed turn up to 3 times with exponential backoff.
  5. Applies an ffmpeg telephone bandpass filter (300Hz–3.4kHz) to correspondent voices.
  6. Keeps the host (Paul) audio clean and unfiltered.
  7. Concatenates all segments in original script order into a single WAV.

Dependencies

  • google-genai (>= 2.0.0)
  • ffmpeg (system)

API Details

Uses the Interactions API with single-speaker TTS — no multi-speaker workaround needed.

Voice Assignment

Voices are assigned dynamically based on [Male] / [Female] gender tags in the script:

SpeakerVoiceAudio Treatment
Paul (host)PuckClean — no filter
Caller [Female] — 1stKoreTelephone filter
Caller [Female] — 2ndAoedeTelephone filter
Caller [Male] — 1stCharonTelephone filter
Caller [Male] — 2ndFenrirTelephone filter

Voices cycle round-robin if there are more callers than available voices. Accent tags ([Accent: Irish], etc.) are injected into the TTS prompt to influence pronunciation.

Telephone Filter

Applied via ffmpeg to correspondent audio segments:

highpass=f=300, lowpass=f=3400, acompressor, volume=1.5

This simulates the standard telephone bandwidth (300Hz–3.4kHz) and adds compression to mimic phone codec dynamics.

Output

  • Primary output: {workspace}/audio/speech/speech.wav
  • Format: WAV, 24kHz, 16-bit PCM, mono
  • Intermediate segments: {workspace}/audio/speech/segments/turn_*.wav

来自 google-gemini 的更多技能

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
使用Gemini API CLI工具的指南。当你需要通过命令行与Gemini API交互、管理代理或生成媒体(图像、……)时使用。
behavioral-evals
google-gemini
创建、运行、修复和推广行为评估的指南。用于验证代理决策逻辑、调试故障、调试提示…
gemini-live-api-dev
google-gemini
通过WebSocket与Gemini进行实时双向流式传输,支持音频、视频和文本对话。具备音频输入/输出(16 kHz PCM)、视频帧、文本以及带语音活动检测的自动转录功能,可实现中断处理。包含原生音频特性:情感对话、主动音频和思考模式;支持同步和异步工具调用的函数调用;以及Google搜索接地功能。提供会话管理,包括上下文压缩、恢复等。
gemini-omni-flash-api
google-gemini
使用此技能进行生成式视频编辑、文本转视频、图像参考视频生成以及首帧到视频的过渡动画,基于…
gemini-api-dev
google-gemini
使用Google的Gemini模型构建应用程序,支持多模态内容、函数调用和结构化输出,覆盖Python、JavaScript、Go和Java语言。可访问当前Gemini 3模型(Pro、Flash、Pro Image),具备100万token上下文;旧版Gemini 2.x和1.5模型已弃用。支持文本生成、图像/音频/视频理解、函数调用、结构化JSON输出、代码执行、上下文缓存和嵌入功能。提供官方SDK:google-genai(Python)...
gemini-interactions-api
google-gemini
Gemini模型与代理的统一接口,支持服务端状态、流式传输和工具编排。兼容当前多款模型(gemini-3-flash-preview、gemini-3-pro-preview、gemini-2.5-flash/pro)及Deep Research代理;自动将已弃用的模型ID替换为当前替代方案。通过previous_interaction_id将对话历史卸载至服务端,实现有状态的多轮交互,无需手动管理历史记录。内置工具编排功能包括...
deliver
google-gemini
将简报的浓缩版发布到 Google Chat 或 Slack 的传入 webhook,使每日运行自行投递——当没有 webhook 时静默跳过……