tts-generation

作者: google-gemini

將廣播節目腳本轉換為語音音訊,並使用Interactions API為通訊員人聲添加電話效果。

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill tts-generation

TTS Generation

Convert the radio show script into speech audio using the Gemini TTS model via the Interactions API. Apply a telephone bandpass filter to correspondent voices so they sound like phone call-ins, while keeping the host's voice clean studio-quality.

Embedded Script

python3 skills/tts-generation/scripts/generate_tts.py --workspace ./workspace

Arguments

ArgumentDefaultDescription
--workspaceworkspaceRoot workspace directory
--workers8Max parallel TTS worker threads

What it does

  1. Reads the script from {workspace}/data/script.md.
  2. Parses it into individual (speaker, text) turns and assigns voices.
  3. Generates TTS for all turns in parallel using a thread pool (default 8 workers).
  4. Retries each failed turn up to 3 times with exponential backoff.
  5. Applies an ffmpeg telephone bandpass filter (300Hz–3.4kHz) to correspondent voices.
  6. Keeps the host (Paul) audio clean and unfiltered.
  7. Concatenates all segments in original script order into a single WAV.

Dependencies

  • google-genai (>= 2.0.0)
  • ffmpeg (system)

API Details

Uses the Interactions API with single-speaker TTS — no multi-speaker workaround needed.

Voice Assignment

Voices are assigned dynamically based on [Male] / [Female] gender tags in the script:

SpeakerVoiceAudio Treatment
Paul (host)PuckClean — no filter
Caller [Female] — 1stKoreTelephone filter
Caller [Female] — 2ndAoedeTelephone filter
Caller [Male] — 1stCharonTelephone filter
Caller [Male] — 2ndFenrirTelephone filter

Voices cycle round-robin if there are more callers than available voices. Accent tags ([Accent: Irish], etc.) are injected into the TTS prompt to influence pronunciation.

Telephone Filter

Applied via ffmpeg to correspondent audio segments:

highpass=f=300, lowpass=f=3400, acompressor, volume=1.5

This simulates the standard telephone bandwidth (300Hz–3.4kHz) and adds compression to mimic phone codec dynamics.

Output

  • Primary output: {workspace}/audio/speech/speech.wav
  • Format: WAV, 24kHz, 16-bit PCM, mono
  • Intermediate segments: {workspace}/audio/speech/segments/turn_*.wav

來自 google-gemini 的更多技能

greeter
google-gemini
一個友善的問候技能
official
code-reviewer
google-gemini
針對本地變更與遠端拉取請求的自動化程式碼審查,提供涵蓋正確性、可維護性及安全性的結構化分析。支援本地檔案系統變更(包含暫存與未暫存)及遠端 PR(依編號或網址),並自動透過 GitHub CLI 進行檢出。從七個面向分析程式碼:正確性、可維護性、可讀性、效率、安全性、邊界情況處理及測試覆蓋率。可執行選用的前置驗證套件(例如 npm run preflight)以提前發現問題。
official
review-duplication
google-gemini
在程式碼審查期間使用此技能,主動檢查程式碼庫中是否存在重複功能、重複造輪子或未能重複使用現有…
official
reconciliation
google-gemini
將已載入的費用與預先解析的發票資料庫進行比對,標記出金額不符、遺漏發票及商家不匹配等差異…
official
agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
official
async-pr-review
google-gemini
當使用者想要開始非同步的 PR 審查、對 PR 執行背景檢查,或查看先前開始的非同步 PR 狀態時,觸發此技能…
official
ci
google-gemini
專為 Gemini CLI 設計的高效能、快速失敗的專業技能
official
critique
google-gemini
專長於審計和修復儲存庫腳本及 GitHub Actions 工作流程,以確保技術穩健性與安全性。
official