tts-generation

โดย google-gemini

แปลงสคริปต์รายการวิทยุเป็นเสียงพูดพร้อมเอฟเฟกต์โทรศัพท์บนเสียงผู้สื่อข่าวโดยใช้ Interactions API

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill tts-generation

TTS Generation

Convert the radio show script into speech audio using the Gemini TTS model via the Interactions API. Apply a telephone bandpass filter to correspondent voices so they sound like phone call-ins, while keeping the host's voice clean studio-quality.

Embedded Script

python3 skills/tts-generation/scripts/generate_tts.py --workspace ./workspace

Arguments

ArgumentDefaultDescription
--workspaceworkspaceRoot workspace directory
--workers8Max parallel TTS worker threads

What it does

  1. Reads the script from {workspace}/data/script.md.
  2. Parses it into individual (speaker, text) turns and assigns voices.
  3. Generates TTS for all turns in parallel using a thread pool (default 8 workers).
  4. Retries each failed turn up to 3 times with exponential backoff.
  5. Applies an ffmpeg telephone bandpass filter (300Hz–3.4kHz) to correspondent voices.
  6. Keeps the host (Paul) audio clean and unfiltered.
  7. Concatenates all segments in original script order into a single WAV.

Dependencies

  • google-genai (>= 2.0.0)
  • ffmpeg (system)

API Details

Uses the Interactions API with single-speaker TTS — no multi-speaker workaround needed.

Voice Assignment

Voices are assigned dynamically based on [Male] / [Female] gender tags in the script:

SpeakerVoiceAudio Treatment
Paul (host)PuckClean — no filter
Caller [Female] — 1stKoreTelephone filter
Caller [Female] — 2ndAoedeTelephone filter
Caller [Male] — 1stCharonTelephone filter
Caller [Male] — 2ndFenrirTelephone filter

Voices cycle round-robin if there are more callers than available voices. Accent tags ([Accent: Irish], etc.) are injected into the TTS prompt to influence pronunciation.

Telephone Filter

Applied via ffmpeg to correspondent audio segments:

highpass=f=300, lowpass=f=3400, acompressor, volume=1.5

This simulates the standard telephone bandwidth (300Hz–3.4kHz) and adds compression to mimic phone codec dynamics.

Output

  • Primary output: {workspace}/audio/speech/speech.wav
  • Format: WAV, 24kHz, 16-bit PCM, mono
  • Intermediate segments: {workspace}/audio/speech/segments/turn_*.wav

Skills เพิ่มเติมจาก google-gemini

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
คู่มือการใช้เครื่องมือ CLI ของ Gemini API ใช้เมื่อคุณต้องการโต้ตอบกับ Gemini API ผ่านทางบรรทัดคำสั่ง จัดการเอเจนต์ หรือสร้างสื่อ (รูปภาพ, …)
behavioral-evals
google-gemini
คำแนะนำสำหรับการสร้าง การรัน การแก้ไข และการส่งเสริมการประเมินพฤติกรรม ใช้เมื่อตรวจสอบตรรกะการตัดสินใจของตัวแทน การแก้ไขข้อบกพร่อง การดีบักพรอมต์…
gemini-live-api-dev
google-gemini
การสตรีมแบบสองทิศทางแบบเรียลไทม์กับ Gemini ผ่าน WebSockets สำหรับการสนทนาด้วยเสียง วีดีโอ และข้อความ รองรับการป้อน/ส่งออกเสียง (16 kHz PCM), เฟรมวีดีโอ, ข้อความ และการถอดความอัตโนมัติพร้อมการตรวจจับกิจกรรมเสียงเพื่อจัดการการขัดจังหวะ รวมถึงคุณสมบัติเสียงแบบเนทีฟ: การสนทนาที่มีอารมณ์, เสียงเชิงรุก และโหมดการคิด; การเรียกใช้ฟังก์ชันสำหรับการใช้เครื่องมือแบบซิงโครนัสและอะซิงโครนัส; และการอ้างอิง Google Search มีการจัดการเซสชันด้วยการบีบอัดบริบท, การกลับมาดำเนินการต่อ และ...
gemini-omni-flash-api
google-gemini
ใช้ทักษะนี้สำหรับการตัดต่อวิดีโอเชิงสร้างสรรค์ การสร้างวิดีโอจากข้อความ การสร้างวิดีโอโดยอ้างอิงจากภาพ และแอนิเมชันเปลี่ยนผ่านจากเฟรมแรกสู่วิดีโอ โดยใช้…
gemini-api-dev
google-gemini
สร้างแอปพลิเคชันด้วยโมเดล Gemini ของ Google รองรับเนื้อหาหลายรูปแบบ การเรียกใช้ฟังก์ชัน และผลลัพธ์ที่มีโครงสร้างใน Python, JavaScript, Go และ Java เข้าถึงโมเดล Gemini 3 ปัจจุบัน (Pro, Flash, Pro Image) พร้อมบริบท 1M โทเค็น; โมเดล Gemini 2.x และ 1.5 รุ่นเก่าถูกเลิกใช้งานแล้ว รองรับการสร้างข้อความ การทำความเข้าใจรูปภาพ/เสียง/วิดีโอ การเรียกใช้ฟังก์ชัน ผลลัพธ์ JSON ที่มีโครงสร้าง การเรียกใช้โค้ด การแคชบริบท และการฝังเวกเตอร์ SDK อย่างเป็นทางการ: google-genai (Python),...
gemini-interactions-api
google-gemini
อินเทอร์เฟซแบบรวมสำหรับโมเดล Gemini และเอเจนต์ พร้อมสถานะฝั่งเซิร์ฟเวอร์ การสตรีม และการจัดระเบียบเครื่องมือ รองรับโมเดลปัจจุบันหลายรุ่น (gemini-3-flash-preview, gemini-3-pro-preview, gemini-2.5-flash/pro) และเอเจนต์ Deep Research; แทนที่รหัสโมเดลที่เลิกใช้งานโดยอัตโนมัติด้วยทางเลือกปัจจุบัน ถ่ายโอนประวัติการสนทนาไปยังเซิร์ฟเวอร์ผ่าน previous_interaction_id สำหรับการโต้ตอบแบบหลายเทิร์นที่มีสถานะโดยไม่ต้องจัดการประวัติด้วยตนเอง มีการจัดระเบียบเครื่องมือในตัวรวมถึง...
deliver
google-gemini
โพสต์เวอร์ชันย่อของสรุปข้อมูลไปยัง Google Chat หรือ Slack incoming webhook เพื่อให้การส่งรายงานประจำวันเกิดขึ้นได้เอง — จะข้ามอย่างเงียบ ๆ เมื่อไม่มี webhook…