tts-generation

โดย google-gemini

Convert radio show script to speech audio with telephone effect on correspondent voices using the Interactions API.

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill tts-generation

TTS Generation

Convert the radio show script into speech audio using the Gemini TTS model via the Interactions API. Apply a telephone bandpass filter to correspondent voices so they sound like phone call-ins, while keeping the host's voice clean studio-quality.

Embedded Script

python3 skills/tts-generation/scripts/generate_tts.py --workspace ./workspace

Arguments

ArgumentDefaultDescription
--workspaceworkspaceRoot workspace directory
--workers8Max parallel TTS worker threads

What it does

  1. Reads the script from {workspace}/data/script.md.
  2. Parses it into individual (speaker, text) turns and assigns voices.
  3. Generates TTS for all turns in parallel using a thread pool (default 8 workers).
  4. Retries each failed turn up to 3 times with exponential backoff.
  5. Applies an ffmpeg telephone bandpass filter (300Hz–3.4kHz) to correspondent voices.
  6. Keeps the host (Paul) audio clean and unfiltered.
  7. Concatenates all segments in original script order into a single WAV.

Dependencies

  • google-genai (>= 2.0.0)
  • ffmpeg (system)

API Details

Uses the Interactions API with single-speaker TTS — no multi-speaker workaround needed.

Voice Assignment

Voices are assigned dynamically based on [Male] / [Female] gender tags in the script:

SpeakerVoiceAudio Treatment
Paul (host)PuckClean — no filter
Caller [Female] — 1stKoreTelephone filter
Caller [Female] — 2ndAoedeTelephone filter
Caller [Male] — 1stCharonTelephone filter
Caller [Male] — 2ndFenrirTelephone filter

Voices cycle round-robin if there are more callers than available voices. Accent tags ([Accent: Irish], etc.) are injected into the TTS prompt to influence pronunciation.

Telephone Filter

Applied via ffmpeg to correspondent audio segments:

highpass=f=300, lowpass=f=3400, acompressor, volume=1.5

This simulates the standard telephone bandwidth (300Hz–3.4kHz) and adds compression to mimic phone codec dynamics.

Output

  • Primary output: {workspace}/audio/speech/speech.wav
  • Format: WAV, 24kHz, 16-bit PCM, mono
  • Intermediate segments: {workspace}/audio/speech/segments/turn_*.wav

Skills เพิ่มเติมจาก google-gemini

greeter
google-gemini
ทักษะการทักทายที่เป็นมิตร
official
code-reviewer
google-gemini
การตรวจสอบโค้ดอัตโนมัติสำหรับการเปลี่ยนแปลงในเครื่องและคำขอดึงข้อมูลระยะไกล พร้อมการวิเคราะห์เชิงโครงสร้างในด้านความถูกต้อง การบำรุงรักษา และความปลอดภัย รองรับทั้งการเปลี่ยนแปลงในระบบไฟล์ท้องถิ่น (ที่จัดเตรียมและไม่ได้จัดเตรียม) และ PR ระยะไกล (ตามหมายเลขหรือ URL) พร้อมการเช็คเอาต์อัตโนมัติผ่าน GitHub CLI วิเคราะห์โค้ดในเจ็ดมิติ: ความถูกต้อง การบำรุงรักษา ความสามารถในการอ่าน ประสิทธิภาพ ความปลอดภัย การจัดการกรณีขอบ และความครอบคลุมของการทดสอบ รันชุดตรวจสอบก่อนการทำงานแบบเลือกได้ (เช่น npm run preflight) เพื่อตรวจจับ...
official
review-duplication
google-gemini
ใช้ทักษะนี้ในระหว่างการตรวจสอบโค้ดเพื่อตรวจสอบฐานโค้ดอย่างเชิงรุกหาฟังก์ชันการทำงานที่ซ้ำซ้อน การสร้างล้อขึ้นมาใหม่ หรือการไม่นำสิ่งที่มีอยู่แล้วกลับมาใช้ใหม่…
official
reconciliation
google-gemini
กระทบยอดค่าใช้จ่ายที่โหลดแล้วกับฐานข้อมูลใบแจ้งหนี้ที่แยกวิเคราะห์ไว้ล่วงหน้า โดยระบุความไม่สอดคล้อง เช่น จำนวนเงินไม่ตรงกัน ใบแจ้งหนี้ที่ขาดหายไป และชื่อผู้ค้าไม่ตรงกัน…
official
agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
official
async-pr-review
google-gemini
เรียกใช้ทักษะนี้เมื่อผู้ใช้ต้องการเริ่มการตรวจสอบ PR แบบอะซิงโครนัส เรียกใช้การตรวจสอบพื้นหลังบน PR หรือตรวจสอบสถานะของ async PR ที่เริ่มไว้ก่อนหน้านี้…
official
ci
google-gemini
ทักษะเฉพาะสำหรับ Gemini CLI ที่ให้ประสิทธิภาพสูงและล้มเหลวเร็ว
official
critique
google-gemini
ความเชี่ยวชาญในการตรวจสอบและแก้ไขสคริปต์ในคลังข้อมูลและเวิร์กโฟลว์ของ GitHub Actions เพื่อให้มั่นใจถึงความแข็งแกร่งทางเทคนิคและความปลอดภัย
official