metadata-generation

作成者: google-gemini

音声と文字起こしをGeminiに送信し、タイムコード付き文字起こし、タイトル、再生時間、要約、日付を含むショーのメタデータ(JSON)を生成します。

npx skills add https://github.com/google-gemini/gemini-managed-agents-templates --skill metadata-generation

Metadata Generation

Generate a structured JSON file containing metadata for the radio show. This is done by sending the final mixed audio file and the original transcript back to Gemini.

Prerequisites

pip install google-genai

Workflow

  1. Read Inputs: Load the transcript from ./workspace/data/script.md and the final audio from ./workspace/audio/final/ai_radio.mp3.
  2. Upload Audio: Upload the audio file to the Gemini Files API.
  3. Generate Metadata: Call Gemini (e.g., gemini-3-flash-preview) using the Interactions API, passing the uploaded audio file and the transcript.
  4. Save Output: Save the resulting JSON to ./workspace/data/show_notes.json.

Python Code

The skill is implemented in scripts/generate_metadata.py.

# Example usage
python3 skills/metadata-generation/scripts/generate_metadata.py --workspace ./workspace

Output

The skill produces a JSON file at ./workspace/data/show_notes.json with the following structure:

{
  "show_title": "...",
  "show_duration": "...",
  "two_sentence_summary": "...",
  "date_of_generation": "YYYY-MM-DD",
  "timecoded_transcript": [
    {
      "timecode": "MM:SS",
      "speaker": "...",
      "text": "..."
    },
    ...
  ]
}

Error Handling

  • If the audio file is missing, the script will fail.
  • If the transcript is missing, it will fail.
  • If the Gemini API call fails, it will report the error and exit.

google-geminiのその他のスキル

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
Gemini API CLIツールの使用ガイド。コマンドラインからGemini APIとやり取りする必要がある場合、エージェントを管理する場合、またはメディア(画像など)を生成する場合に使用します。
behavioral-evals
google-gemini
行動評価の作成、実行、修正、促進に関するガイダンス。エージェントの意思決定ロジックの検証、障害のデバッグ、プロンプトのデバッグなどに使用します。
gemini-live-api-dev
google-gemini
Geminiを介したWebSocket上のリアルタイム双方向ストリーミングにより、音声、動画、テキストの会話を実現。音声入出力(16kHz PCM)、動画フレーム、テキスト、および割り込み処理のための音声アクティビティ検出による自動文字起こしをサポート。ネイティブ音声機能(感情対話、プロアクティブ音声、思考モード)、同期・非同期ツール使用のための関数呼び出し、Google Searchグラウンディングを搭載。コンテキスト圧縮、再開などを備えたセッション管理を提供。
gemini-omni-flash-api
google-gemini
このスキルは、生成型ビデオ編集、テキストからビデオ、画像参照によるビデオ生成、および最初のフレームからビデオへのトランジションアニメーションに使用します…
gemini-api-dev
google-gemini
GoogleのGeminiモデルを使用してアプリケーションを構築します。マルチモーダルコンテンツ、関数呼び出し、構造化出力をサポートし、Python、JavaScript、Go、Javaに対応しています。現在のGemini 3モデル(Pro、Flash、Pro Image)にアクセス可能で、100万トークンのコンテキストを備えています。レガシーのGemini 2.xおよび1.5モデルは非推奨です。テキスト生成、画像/音声/動画の理解、関数呼び出し、構造化JSON出力、コード実行、コンテキストキャッシング、埋め込みをサポートしています。公式SDKとしてgoogle-genai(Python)などが利用可能です。
gemini-interactions-api
google-gemini
Geminiモデルとエージェントのための統合インターフェース。サーバーサイドの状態、ストリーミング、ツールオーケストレーションを備えています。複数の現行モデル(gemini-3-flash-preview、gemini-3-pro-preview、gemini-2.5-flash/pro)およびDeep Researchエージェントをサポート。非推奨のモデルIDを現行の代替モデルに自動的に置き換えます。previous_interaction_idを介して会話履歴をサーバーにオフロードし、手動での履歴管理なしでステートフルなマルチターン対話を実現。組み込みのツールオーケストレーションを含む...
deliver
google-gemini
ブリーフィングの要約版をGoogle ChatまたはSlackの受信ウェブフックに投稿するので、毎日の実行が自動的に配信される — ウェブフックが設定されていない場合は静かにスキップする…