audio-mixing

作者: google-gemini

將語音音訊與背景音樂混合,製作出精緻的廣播節目檔案。

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill audio-mixing

Audio Mixing

Combine the TTS speech audio and Lyria background music into a single, polished radio show file.

Embedded Script

python3 skills/audio-mixing/scripts/mix_audio.py --workspace ./workspace

Arguments

ArgumentDefaultDescription
--workspaceworkspaceRoot workspace directory

What it does

  1. Loads speech from {workspace}/audio/speech/speech.wav.
  2. Adds 3 seconds of silence padding to the end of the speech to prevent it from being faded out.
  3. Loads background music from {workspace}/audio/music/background.mp3 (if exists).
  4. Loops music to match speech duration, lowers volume to -18dB.
  5. Overlays speech on music with a 1-second music intro.
  6. Adds fade-in/fade-out.
  7. Exports as MP3.

Dependencies

  • pydub
  • ffmpeg (system)

Mixing Guidelines

ElementLevelNotes
Speech0 dBUntouched, full volume
Background music-18 dBBarely audible — subtle bed under speech

Transitions

  • Music fade-in: 3 seconds
  • Music fade-out: 5 seconds
  • Overall fade-in: 500ms
  • Overall fade-out: 2 seconds

Output

FilePathFormat
MP3 (distribution){workspace}/audio/final/ai_radio.mp3MP3, 192kbps

Fallback

If no background music exists, produces speech-only output with just fade-in/fade-out applied.

來自 google-gemini 的更多技能

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
使用 Gemini API CLI 工具的指南。當你需要透過命令列與 Gemini API 互動、管理代理或生成媒體(圖片、……)時使用。
behavioral-evals
google-gemini
建立、執行、修正及推廣行為評估的指引。用於驗證代理決策邏輯、除錯失敗、除錯提示…
gemini-live-api-dev
google-gemini
通過WebSocket與Gemini進行即時雙向串流,支援音訊、視訊和文字對話。支援音訊輸入/輸出(16 kHz PCM)、視訊幀、文字,以及具備語音活動偵測的自動轉錄功能,可處理中斷情況。包含原生音訊功能:情感對話、主動音訊和思考模式;支援同步和非同步工具使用的函式呼叫;以及Google Search基礎驗證。提供具備上下文壓縮、恢復功能的會話管理,以及...
gemini-omni-flash-api
google-gemini
使用此技能進行生成式影片編輯、文字轉影片、圖片參考影片生成,以及基於…的首幀轉場動畫。
gemini-api-dev
google-gemini
We need to translate the given text from English to Traditional Chinese. The text describes building applications with Google's Gemini models, mentioning multimodal content, function calling, structured outputs, supported languages, model versions, features, and SDKs. We must preserve the name "gemini-api-dev" but it's not in the text, so ignore. Also preserve technical terms like "Gemini", "Pro", "Flash", "Pro Image", "1M token context", "Gemini 2.x", "1.5", "JSON", "SDKs", "google-genai", etc. No extra commentary. Output only the translation. Translation: 使用 Google 的 Gemini 模型建置應用程式,支援多模態內容、函式呼叫與結構化輸出,涵蓋 Python、JavaScript、Go 及 Java。可存取最新的 Gemini 3 模型(Pro、Flash、Pro Image),具備 100 萬 Token 上下文;舊版 Gemini 2.x 與 1.5 模型已棄用。支援文字生成、圖片/音訊
gemini-interactions-api
google-gemini
Gemini模型與代理的統一介面,具備伺服器端狀態、串流與工具編排功能。支援多種當前模型(gemini-3-flash-preview、gemini-3-pro-preview、gemini-2.5-flash/pro)及Deep Research代理;自動將已棄用的模型ID替換為當前替代方案。透過previous_interaction_id將對話歷史卸載至伺服器,實現有狀態的多輪互動,無需手動管理歷史記錄。內建工具編排功能,包括...
deliver
google-gemini
將簡報的精簡版本發布到 Google Chat 或 Slack 的 incoming webhook,讓每日執行自動送達——若未設定 webhook 則靜默跳過…