metadata-generation

द्वारा google-gemini

ऑडियो और ट्रांसक्रिप्ट को Gemini पर भेजकर टाइमकोडेड ट्रांसक्रिप्ट, शीर्षक, अवधि, सारांश और तारीख वाला शो मेटाडेटा (JSON) उत्पन्न करें।

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill metadata-generation

Metadata Generation

Generate a structured JSON file containing metadata for the radio show. This is done by sending the final mixed audio file and the original transcript back to Gemini.

Prerequisites

pip install google-genai

Workflow

  1. Read Inputs: Load the transcript from ./workspace/data/script.md and the final audio from ./workspace/audio/final/ai_radio.mp3.
  2. Upload Audio: Upload the audio file to the Gemini Files API.
  3. Generate Metadata: Call Gemini (e.g., gemini-3-flash-preview) using the Interactions API, passing the uploaded audio file and the transcript.
  4. Save Output: Save the resulting JSON to ./workspace/data/show_notes.json.

Python Code

The skill is implemented in scripts/generate_metadata.py.

# Example usage
python3 skills/metadata-generation/scripts/generate_metadata.py --workspace ./workspace

Output

The skill produces a JSON file at ./workspace/data/show_notes.json with the following structure:

{
  "show_title": "...",
  "show_duration": "...",
  "two_sentence_summary": "...",
  "date_of_generation": "YYYY-MM-DD",
  "timecoded_transcript": [
    {
      "timecode": "MM:SS",
      "speaker": "...",
      "text": "..."
    },
    ...
  ]
}

Error Handling

  • If the audio file is missing, the script will fail.
  • If the transcript is missing, it will fail.
  • If the Gemini API call fails, it will report the error and exit.

google-gemini की और Skills

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
Gemini API CLI टूल का उपयोग करने के लिए मार्गदर्शिका। इसका उपयोग तब करें जब आपको कमांड लाइन के माध्यम से Gemini API से संवाद करना हो, एजेंट प्रबंधित करने हों, या मीडिया (चित्र, …) उत्पन्न करना हो।
behavioral-evals
google-gemini
व्यवहारिक मूल्यांकन बनाने, चलाने, ठीक करने और बढ़ावा देने के लिए मार्गदर्शन। एजेंट निर्णय तर्क को सत्यापित करने, विफलताओं को डीबग करने, प्रॉम्प्ट को डीबग करने के दौरान उपयोग करें…
gemini-live-api-dev
google-gemini
जेमिनी के साथ वेबसॉकेट पर ऑडियो, वीडियो और टेक्स्ट वार्तालापों के लिए रीयल-टाइम द्विदिश स्ट्रीमिंग। ऑडियो इनपुट/आउटपुट (16 kHz PCM), वीडियो फ्रेम, टेक्स्ट और व्यवधान प्रबंधन के लिए वॉयस एक्टिविटी डिटेक्शन के साथ स्वचालित ट्रांसक्रिप्शन का समर्थन करता है। इसमें मूल ऑडियो सुविधाएँ शामिल हैं: भावात्मक संवाद, सक्रिय ऑडियो और विचार मोड; सिंक्रोनस और एसिंक्रोनस
gemini-omni-flash-api
google-gemini
जनरेटिव वीडियो संपादन, टेक्स्ट-टू-वीडियो, इमेज-रेफरेंस्ड वीडियो जनरेशन, और फर्स्ट-फ्रेम-टू-वीडियो ट्रांज़िशन एनिमेशन के लिए इस स्किल का उपयोग करें...
gemini-api-dev
google-gemini
Google के Gemini मॉडल के साथ एप्लिकेशन बनाएं, जो Python, JavaScript, Go और Java में मल्टीमॉडल कंटेंट, फंक्शन कॉलिंग और स्ट्रक्चर्ड आउटपुट को सपोर्ट करता है। 1M टोकन कॉन्टेक्स्ट के साथ वर्तमान Gemini 3 मॉडल (Pro, Flash, Pro Image) तक पहुंचें; लीगेसी Gemini 2.x और 1.5 मॉडल डिप्रीकेटेड हैं। टेक्स्ट जनरेशन, इमेज/ऑडियो/वीडियो समझ, फंक्शन कॉलिंग, स्ट्रक्चर्ड JSON आउटपुट, कोड एक्जी
gemini-interactions-api
google-gemini
Gemini मॉडल और एजेंटों के लिए एकीकृत इंटरफ़ेस, जिसमें सर्वर-साइड स्थिति, स्ट्रीमिंग और टूल ऑर्केस्ट्रेशन शामिल है। कई वर्तमान मॉडलों (gemini-3-flash-preview, gemini-3-pro-preview, gemini-2.5-flash/pro) और डीप रिसर्च एजेंट का समर्थन करता है; पुराने मॉडल आईडी को स्वचालित रूप से वर्तमान विकल्पों से बदल देता है। मैन्युअल इतिहास प्रबंधन के बिना स्टेटफुल मल्टी-टर्न इंटरैक्शन के लिए previous_interaction_id के माध्य
deliver
google-gemini
ब्रिफिंग का संक्षिप्त संस्करण Google Chat या Slack इनकमिंग वेबहुक पर पोस्ट करता है, ताकि दैनिक रन स्वयं वितरित हो जाए — जब कोई वेबहुक न हो तो चुपचाप छोड़ देता है…