tts-generation

Converta o script do programa de rádio em áudio de fala com efeito telefônico nas vozes dos correspondentes usando a API Interactions.

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill tts-generation

TTS Generation

Convert the radio show script into speech audio using the Gemini TTS model via the Interactions API. Apply a telephone bandpass filter to correspondent voices so they sound like phone call-ins, while keeping the host's voice clean studio-quality.

Embedded Script

python3 skills/tts-generation/scripts/generate_tts.py --workspace ./workspace

Arguments

ArgumentDefaultDescription
--workspaceworkspaceRoot workspace directory
--workers8Max parallel TTS worker threads

What it does

  1. Reads the script from {workspace}/data/script.md.
  2. Parses it into individual (speaker, text) turns and assigns voices.
  3. Generates TTS for all turns in parallel using a thread pool (default 8 workers).
  4. Retries each failed turn up to 3 times with exponential backoff.
  5. Applies an ffmpeg telephone bandpass filter (300Hz–3.4kHz) to correspondent voices.
  6. Keeps the host (Paul) audio clean and unfiltered.
  7. Concatenates all segments in original script order into a single WAV.

Dependencies

  • google-genai (>= 2.0.0)
  • ffmpeg (system)

API Details

Uses the Interactions API with single-speaker TTS — no multi-speaker workaround needed.

Voice Assignment

Voices are assigned dynamically based on [Male] / [Female] gender tags in the script:

SpeakerVoiceAudio Treatment
Paul (host)PuckClean — no filter
Caller [Female] — 1stKoreTelephone filter
Caller [Female] — 2ndAoedeTelephone filter
Caller [Male] — 1stCharonTelephone filter
Caller [Male] — 2ndFenrirTelephone filter

Voices cycle round-robin if there are more callers than available voices. Accent tags ([Accent: Irish], etc.) are injected into the TTS prompt to influence pronunciation.

Telephone Filter

Applied via ffmpeg to correspondent audio segments:

highpass=f=300, lowpass=f=3400, acompressor, volume=1.5

This simulates the standard telephone bandwidth (300Hz–3.4kHz) and adds compression to mimic phone codec dynamics.

Output

  • Primary output: {workspace}/audio/speech/speech.wav
  • Format: WAV, 24kHz, 16-bit PCM, mono
  • Intermediate segments: {workspace}/audio/speech/segments/turn_*.wav

Mais skills de google-gemini

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
Guia para usar a ferramenta de linha de comando da API Gemini. Use quando precisar interagir com a API Gemini via linha de comando, gerenciar agentes ou gerar mídia (imagens,…
behavioral-evals
google-gemini
Orientação para criar, executar, corrigir e promover avaliações comportamentais. Use ao verificar a lógica de decisão do agente, depurar falhas, depurar prompt…
gemini-live-api-dev
google-gemini
We need to translate the given text from English to Brazilian Portuguese. The text describes a real-time bidirectional streaming capability with Gemini over WebSockets. It mentions audio, video, text, audio input/output specs, video frames, automatic transcriptions, voice activity detection, interruption handling, native audio features (affective dialog, proactive audio, thinking mode), function calling, Google Search grounding, session management with context compression, resumption, etc. We must preserve product names, protocol names, URLs, numbers, technical terms. The name "gemini-live-api-dev" is not in the text, so we don't include it. We only translate the text inside <text>. No extra commentary, no labels. The text ends with "and..." so we keep that as is. Translation: "Streaming bidirecional em tempo real com Gemini via WebSockets para conversas de áudio, vídeo e texto. Suporta entrada/saída de áudio (PCM 16 kHz), quadros de vídeo, texto e transcrições automáticas com detecção
gemini-omni-flash-api
google-gemini
Use esta habilidade para edição generativa de vídeo, texto para vídeo, geração de vídeo com referência de imagem e animações de transição de primeiro quadro para vídeo usando o…
gemini-api-dev
google-gemini
We need to translate the given text from English to Brazilian Portuguese. The text describes building applications with Google's Gemini models. We must preserve product names, protocol names, URLs, numbers, technical terms. The name "gemini-api-dev" is not in the text, so we don't include it. We translate only the text inside <text>. No extra commentary, no labels. The text: "Build applications with Google's Gemini models, supporting multimodal content, function calling, and structured outputs across Python, JavaScript, Go, and Java. Access current Gemini 3 models (Pro, Flash, Pro Image) with 1M token context; legacy Gemini 2.x and 1.5 models are deprecated Supports text generation, image/audio/video understanding, function calling, structured JSON output, code execution, context caching, and embeddings Official SDKs available: google-genai (Python),..." We need to translate fluently but preserve terms like "Gemini models", "multimodal content", "function calling", "structured outputs", "Python", "JavaScript", "Go",
gemini-interactions-api
google-gemini
Interface unificado para modelos e agentes Gemini com estado no servidor, streaming e orquestração de ferramentas. Suporta vários modelos atuais (gemini-3-flash-preview, gemini-3-pro-preview, gemini-2.5-flash/pro) e o agente Deep Research; substitui automaticamente IDs de modelos obsoletos por alternativas atuais. Descarrega o histórico de conversas para o servidor via previous_interaction_id para interações multi-turno com estado, sem gerenciamento manual de histórico. Orquestração de ferramentas integrada, incluindo...
deliver
google-gemini
Publica uma versão condensada do briefing em um webhook de entrada do Google Chat ou Slack, para que a execução diária se entregue sozinha — pula silenciosamente quando não há webhook…