tts-generation

작성자: google-gemini

라디오 쇼 대본을 음성 오디오로 변환하고, Interactions API를 사용하여 특파원 목소리에 전화 효과를 적용합니다.

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill tts-generation

TTS Generation

Convert the radio show script into speech audio using the Gemini TTS model via the Interactions API. Apply a telephone bandpass filter to correspondent voices so they sound like phone call-ins, while keeping the host's voice clean studio-quality.

Embedded Script

python3 skills/tts-generation/scripts/generate_tts.py --workspace ./workspace

Arguments

ArgumentDefaultDescription
--workspaceworkspaceRoot workspace directory
--workers8Max parallel TTS worker threads

What it does

  1. Reads the script from {workspace}/data/script.md.
  2. Parses it into individual (speaker, text) turns and assigns voices.
  3. Generates TTS for all turns in parallel using a thread pool (default 8 workers).
  4. Retries each failed turn up to 3 times with exponential backoff.
  5. Applies an ffmpeg telephone bandpass filter (300Hz–3.4kHz) to correspondent voices.
  6. Keeps the host (Paul) audio clean and unfiltered.
  7. Concatenates all segments in original script order into a single WAV.

Dependencies

  • google-genai (>= 2.0.0)
  • ffmpeg (system)

API Details

Uses the Interactions API with single-speaker TTS — no multi-speaker workaround needed.

Voice Assignment

Voices are assigned dynamically based on [Male] / [Female] gender tags in the script:

SpeakerVoiceAudio Treatment
Paul (host)PuckClean — no filter
Caller [Female] — 1stKoreTelephone filter
Caller [Female] — 2ndAoedeTelephone filter
Caller [Male] — 1stCharonTelephone filter
Caller [Male] — 2ndFenrirTelephone filter

Voices cycle round-robin if there are more callers than available voices. Accent tags ([Accent: Irish], etc.) are injected into the TTS prompt to influence pronunciation.

Telephone Filter

Applied via ffmpeg to correspondent audio segments:

highpass=f=300, lowpass=f=3400, acompressor, volume=1.5

This simulates the standard telephone bandwidth (300Hz–3.4kHz) and adds compression to mimic phone codec dynamics.

Output

  • Primary output: {workspace}/audio/speech/speech.wav
  • Format: WAV, 24kHz, 16-bit PCM, mono
  • Intermediate segments: {workspace}/audio/speech/segments/turn_*.wav

google-gemini의 다른 스킬

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
Gemini API CLI 도구 사용 가이드입니다. 명령줄을 통해 Gemini API와 상호작용하거나, 에이전트를 관리하거나, 미디어(이미지 등)를 생성해야 할 때 사용하세요.
behavioral-evals
google-gemini
행동 평가를 생성, 실행, 수정 및 홍보하기 위한 지침입니다. 에이전트 결정 로직 검증, 오류 디버깅, 프롬프트 디버깅 등에 사용하세요.
gemini-live-api-dev
google-gemini
WebSocket을 통해 Gemini와 실시간 양방향 스트리밍을 지원하여 오디오, 비디오, 텍스트 대화를 처리합니다. 오디오 입력/출력(16kHz PCM), 비디오 프레임, 텍스트, 음성 활동 감지를 통한 자동 전사 및 인터럽트 처리를 지원합니다. 네이티브 오디오 기능(감정 대화, 능동적 오디오, 사고 모드), 동기 및 비동기 도구 사용을 위한 함수 호출, Google Search 접지 기능을 포함합니다. 컨텍스트 압축, 재개 등을 통한 세션 관리를 제공합니다.
gemini-omni-flash-api
google-gemini
이 스킬을 사용하여 생성형 비디오 편집, 텍스트-투-비디오, 이미지 참조 비디오 생성, 첫 프레임-투-비디오 전환 애니메이션 등을 수행할 수 있습니다…
gemini-api-dev
google-gemini
Google의 Gemini 모델로 애플리케이션을 구축하며, 멀티모달 콘텐츠, 함수 호출, 구조화된 출력을 Python, JavaScript, Go, Java에서 지원합니다. 최신 Gemini 3 모델(Pro, Flash, Pro Image)에 1M 토큰 컨텍스트로 접근 가능하며, 레거시 Gemini 2.x 및 1.5 모델은 지원 중단되었습니다. 텍스트 생성, 이미지/오디오/비디오 이해, 함수 호출, 구조화된 JSON 출력, 코드 실행, 컨텍스트 캐싱, 임베딩을 지원합니다. 공식 SDK 제공: google-genai (Python),...
gemini-interactions-api
google-gemini
Gemini 모델 및 에이전트를 위한 통합 인터페이스로, 서버 측 상태, 스트리밍 및 도구 오케스트레이션을 제공합니다. 여러 현재 모델(gemini-3-flash-preview, gemini-3-pro-preview, gemini-2.5-flash/pro)과 Deep Research 에이전트를 지원하며, 더 이상 사용되지 않는 모델 ID를 현재 대안으로 자동 대체합니다. previous_interaction_id를 통해 대화 기록을 서버에 오프로드하여 수동 기록 관리 없이 상태 저장 다중 턴 상호작용을 가능하게 합니다. 내장된 도구 오케스트레이션을 포함합니다...
deliver
google-gemini
브리핑의 축약 버전을 Google Chat 또는 Slack 인커밍 웹훅에 게시하여 일일 실행이 자동으로 전달되게 합니다 — 웹훅이 없으면 조용히 건너뜁니다…