audio-mixing

द्वारा google-gemini

भाषण ऑडियो और पृष्ठभूमि संगीत को मिलाकर एक पॉलिश रेडियो शो फ़ाइल बनाएं।

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill audio-mixing

Audio Mixing

Combine the TTS speech audio and Lyria background music into a single, polished radio show file.

Embedded Script

python3 skills/audio-mixing/scripts/mix_audio.py --workspace ./workspace

Arguments

ArgumentDefaultDescription
--workspaceworkspaceRoot workspace directory

What it does

  1. Loads speech from {workspace}/audio/speech/speech.wav.
  2. Adds 3 seconds of silence padding to the end of the speech to prevent it from being faded out.
  3. Loads background music from {workspace}/audio/music/background.mp3 (if exists).
  4. Loops music to match speech duration, lowers volume to -18dB.
  5. Overlays speech on music with a 1-second music intro.
  6. Adds fade-in/fade-out.
  7. Exports as MP3.

Dependencies

  • pydub
  • ffmpeg (system)

Mixing Guidelines

ElementLevelNotes
Speech0 dBUntouched, full volume
Background music-18 dBBarely audible — subtle bed under speech

Transitions

  • Music fade-in: 3 seconds
  • Music fade-out: 5 seconds
  • Overall fade-in: 500ms
  • Overall fade-out: 2 seconds

Output

FilePathFormat
MP3 (distribution){workspace}/audio/final/ai_radio.mp3MP3, 192kbps

Fallback

If no background music exists, produces speech-only output with just fade-in/fade-out applied.

google-gemini की और Skills

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
Gemini API CLI टूल का उपयोग करने के लिए मार्गदर्शिका। इसका उपयोग तब करें जब आपको कमांड लाइन के माध्यम से Gemini API से संवाद करना हो, एजेंट प्रबंधित करने हों, या मीडिया (चित्र, …) उत्पन्न करना हो।
behavioral-evals
google-gemini
व्यवहारिक मूल्यांकन बनाने, चलाने, ठीक करने और बढ़ावा देने के लिए मार्गदर्शन। एजेंट निर्णय तर्क को सत्यापित करने, विफलताओं को डीबग करने, प्रॉम्प्ट को डीबग करने के दौरान उपयोग करें…
gemini-live-api-dev
google-gemini
जेमिनी के साथ वेबसॉकेट पर ऑडियो, वीडियो और टेक्स्ट वार्तालापों के लिए रीयल-टाइम द्विदिश स्ट्रीमिंग। ऑडियो इनपुट/आउटपुट (16 kHz PCM), वीडियो फ्रेम, टेक्स्ट और व्यवधान प्रबंधन के लिए वॉयस एक्टिविटी डिटेक्शन के साथ स्वचालित ट्रांसक्रिप्शन का समर्थन करता है। इसमें मूल ऑडियो सुविधाएँ शामिल हैं: भावात्मक संवाद, सक्रिय ऑडियो और विचार मोड; सिंक्रोनस और एसिंक्रोनस
gemini-omni-flash-api
google-gemini
जनरेटिव वीडियो संपादन, टेक्स्ट-टू-वीडियो, इमेज-रेफरेंस्ड वीडियो जनरेशन, और फर्स्ट-फ्रेम-टू-वीडियो ट्रांज़िशन एनिमेशन के लिए इस स्किल का उपयोग करें...
gemini-api-dev
google-gemini
Google के Gemini मॉडल के साथ एप्लिकेशन बनाएं, जो Python, JavaScript, Go और Java में मल्टीमॉडल कंटेंट, फंक्शन कॉलिंग और स्ट्रक्चर्ड आउटपुट को सपोर्ट करता है। 1M टोकन कॉन्टेक्स्ट के साथ वर्तमान Gemini 3 मॉडल (Pro, Flash, Pro Image) तक पहुंचें; लीगेसी Gemini 2.x और 1.5 मॉडल डिप्रीकेटेड हैं। टेक्स्ट जनरेशन, इमेज/ऑडियो/वीडियो समझ, फंक्शन कॉलिंग, स्ट्रक्चर्ड JSON आउटपुट, कोड एक्जी
gemini-interactions-api
google-gemini
Gemini मॉडल और एजेंटों के लिए एकीकृत इंटरफ़ेस, जिसमें सर्वर-साइड स्थिति, स्ट्रीमिंग और टूल ऑर्केस्ट्रेशन शामिल है। कई वर्तमान मॉडलों (gemini-3-flash-preview, gemini-3-pro-preview, gemini-2.5-flash/pro) और डीप रिसर्च एजेंट का समर्थन करता है; पुराने मॉडल आईडी को स्वचालित रूप से वर्तमान विकल्पों से बदल देता है। मैन्युअल इतिहास प्रबंधन के बिना स्टेटफुल मल्टी-टर्न इंटरैक्शन के लिए previous_interaction_id के माध्य
deliver
google-gemini
ब्रिफिंग का संक्षिप्त संस्करण Google Chat या Slack इनकमिंग वेबहुक पर पोस्ट करता है, ताकि दैनिक रन स्वयं वितरित हो जाए — जब कोई वेबहुक न हो तो चुपचाप छोड़ देता है…