gemini-live-api-dev

We need to translate the given text from English to French, preserving the name "gemini-live-api-dev" if it appears, but it does not appear in the text. The text is a description of an agent skill. We must not add any extra commentary, labels, or formatting. Just the translation. The text: "Real-time bidirectional streaming with Gemini over WebSockets for audio, video, and text conversations. Supports audio input/output (16 kHz PCM), video frames, text, and automatic transcriptions with voice activity detection for interruption handling Includes native audio features: affective dialog, proactive audio, and thinking mode; function calling for synchronous and asynchronous tool use; and Google Search grounding Offers session management with context compression, resumption, and..." We need to translate this into French. Note: "Gemini" is a product name, keep as is. "WebSockets" is a protocol name, keep as is. "16 kHz PCM" is technical, keep as is. "Google Search" is a product name, keep as is. Also "affective dialog",

npx skills add https://github.com/google-gemini/gemini-skills --skill gemini-live-api-dev

Gemini Live API Development Skill

Overview

The Live API enables low-latency, real-time voice and video interactions with Gemini over WebSockets. It processes continuous streams of audio, video, or text to deliver immediate, human-like spoken responses and background reasoning.

Key capabilities:

  • Bidirectional audio streaming — real-time mic-to-speaker conversations
  • Background reasoning (extended thinking) — multi-step background reasoning with spoken conversational fillers
  • Live streaming transcription — real-time speech-to-text with interim and finalized streams
  • Video streaming — send camera/screen frames alongside audio
  • Text input/output — send and receive text within a live session
  • Audio transcriptions — get text transcripts of both input and output audio
  • Voice Activity Detection (VAD) — automatic server VAD, client-side Hybrid VAD, and manual Push-to-Talk
  • Asynchronous function calling — non-blocking tool execution while audio continues streaming
  • Full-session client content — inject and update conversation turns mid-stream
  • Session management — context compression, session resumption, GoAway signals
  • Ephemeral tokens — secure client-side authentication

[!NOTE] The Live API connects directly via WebSockets. For WebRTC support or simplified integration, use a partner integration.

Models

Current Models (Use These)

  • gemini-3.8-live — Default option for most low-latency voice agent experiences and real-time dialogue without reasoning delays. Supports interleaved reasoning, asynchronous function calling by default (behavior: NON_BLOCKING), and full-session client content updates.
  • gemini-3.8-live-extended-thinking — High-reasoning audio-to-audio model recommended when higher background reasoning is required during live interactions. Processes background reasoning and async tool calls (behavior: NON_BLOCKING required) while streaming continuous spoken conversational fillers; lifecycle managed via interaction_status (IN_PROGRESS vs IDLE).
  • gemini-3.5-transcribe-live — Real-time streaming speech-to-text with interim hypotheses, finalized transcripts, smart formatting, and Hybrid VAD.
  • gemini-3.5-live-translate-preview — Real-time speech-to-speech streaming translation across 70+ languages.

[!WARNING] Legacy Models (gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-*, gemini-live-2.5-flash-preview, gemini-2.0-flash-live-001): Read references/migration.md for breaking protocol changes (behavior: "NON_BLOCKING", thinking_level, interaction_status, send_client_content).

SDKs

  • Python: google-genai >= 2.3.0 — pip install -U google-genai
  • JavaScript/TypeScript: @google/genai >= 2.3.0 — npm install @google/genai

[!WARNING] Legacy SDKs google-generativeai (Python) and @google/generative-ai (JS) are deprecated. Never use them.

Partner Integrations

To streamline real-time audio/video app development, use a third-party integration supporting the Gemini Live API over WebRTC or WebSockets:

  • LiveKit — Use the Gemini Live API with LiveKit Agents.
  • Pipecat by Daily — Create a real-time AI chatbot using Gemini Live and Pipecat.
  • Fishjam by Software Mansion — Create live video and audio streaming applications with Fishjam.
  • Vision Agents by Stream — Build real-time voice and video AI applications with Vision Agents.
  • Voximplant — Connect inbound and outbound calls to Live API with Voximplant.
  • Firebase AI SDK — Get started with the Gemini Live API using Firebase AI Logic.

Audio Formats

  • Input: Raw PCM, little-endian, 16-bit, mono. 16kHz native (will resample others). MIME type: audio/pcm;rate=16000
  • Output: Raw PCM, little-endian, 16-bit, mono. 24kHz sample rate.

[!IMPORTANT] Use send_realtime_input / sendRealtimeInput for all real-time streaming user input (audio, video, and text). On Gemini 3.8 models, send_client_content / sendClientContent is supported across the full session lifecycle with explicit roles (user or model) to inject conversation context (turn_complete=true unconditionally interrupts active generation).

[!WARNING] Do not use media in sendRealtimeInput. Use the specific keys: audio for audio data, video for images/video frames, and text for text input.


Quick Start

Authentication

Python

from google import genai

client = genai.Client(api_key="YOUR_API_KEY")

JavaScript

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI({ apiKey: 'YOUR_API_KEY' });

Connecting to the Live API

Python

from google.genai import types

config = types.LiveConnectConfig(
    response_modalities=[types.Modality.AUDIO],
    system_instruction=types.Content(
        parts=[types.Part(text="You are a helpful assistant.")]
    )
)

async with client.aio.live.connect(model="gemini-3.8-live", config=config) as session:
    pass  # Session is active

JavaScript

const session = await ai.live.connect({
  model: 'gemini-3.8-live',
  config: {
    responseModalities: ['audio'],
    systemInstruction: { parts: [{ text: 'You are a helpful assistant.' }] }
  },
  callbacks: {
    onopen: () => console.log('Connected'),
    onmessage: (response) => console.log('Message:', response),
    onerror: (error) => console.error('Error:', error),
    onclose: () => console.log('Closed')
  }
});

Sending Text

Python

await session.send_realtime_input(text="Hello, how are you?")

JavaScript

session.sendRealtimeInput({ text: 'Hello, how are you?' });

Sending Audio

Python

await session.send_realtime_input(
    audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000")
)

JavaScript

session.sendRealtimeInput({
  audio: { data: chunk.toString('base64'), mimeType: 'audio/pcm;rate=16000' }
});

Sending Video

Python

# frame: raw JPEG-encoded bytes
await session.send_realtime_input(
    video=types.Blob(data=frame, mime_type="image/jpeg")
)

JavaScript

session.sendRealtimeInput({
  video: { data: frame.toString('base64'), mimeType: 'image/jpeg' }
});

Receiving Audio and Text

[!IMPORTANT] A single server event can contain multiple content parts simultaneously (e.g., audio chunks and transcript). Always process all parts in each event to avoid missing content.

Python

async for response in session.receive():
    content = response.server_content
    if content:
        # Audio — process ALL parts in each event
        if content.model_turn:
            for part in content.model_turn.parts:
                if part.inline_data:
                    audio_data = part.inline_data.data
        # Transcription
        if content.input_transcription:
            print(f"User: {content.input_transcription.text}")
        if content.output_transcription:
            print(f"Gemini: {content.output_transcription.text}")
        # Interruption
        if content.interrupted is True:
            pass  # Stop playback, clear audio queue

JavaScript

// Inside the onmessage callback
const content = response.serverContent;
if (content?.modelTurn?.parts) {
  for (const part of content.modelTurn.parts) {
    if (part.inlineData) {
      const audioData = part.inlineData.data; // Base64 encoded
    }
  }
}
if (content?.inputTranscription) console.log('User:', content.inputTranscription.text);
if (content?.outputTranscription) console.log('Gemini:', content.outputTranscription.text);
if (content?.interrupted) { /* Stop playback, clear audio queue */ }

Background Reasoning (Extended Thinking)

Use gemini-3.8-live-extended-thinking when your voice agent must evaluate complex data, plan multiple steps, or handle long-running tools. The model speaks natural conversational fillers (e.g. "Checking flight options now...") while executing asynchronous tools in the background.

Key requirements:

  • Thinking config: Set thinking_config=types.ThinkingConfig(thinking_level="low") ("minimal" | "low" | "medium" | "high").
  • Non-blocking tools: All function declarations must set behavior="NON_BLOCKING". Synchronous blocking mode is not supported and returns an error.
  • Lifecycle tracking (interaction_status): Do not rely on turn_complete=True alone to detect turn completion. Monitor message.interaction_status (Python) / message.interactionStatus (JS):
    • "IN_PROGRESS": Server is reasoning, speaking conversational fillers, or waiting for async tool responses.
    • "IDLE": Server has completed all background reasoning and tool calls; session is ready for user input.

See references/migration.md and the Thinking in Live API Guide for complete Python and JavaScript implementation examples.


Live Translation (Gemini Live Translate)

The Live API supports real-time, low-latency streaming translation of speech (audio) across 70+ languages. For full details on options and capabilities, see the Live Translate Guide.

Model

  • gemini-3.5-live-translate-preview — The recommended translation model for all Live Translate use cases.

Configuration (TranslationConfig)

To enable translation, specify a TranslationConfig object inside your live session setup:

  • Python SDK: Configure the connection using translation_config on LiveConnectConfig:
    config = types.LiveConnectConfig(
        response_modalities=[types.Modality.AUDIO],
        translation_config=types.TranslationConfig(
            target_language_code="es",  # Target language code (e.g. es, fr, pl)
            echo_target_language=True,
        ),
        input_audio_transcription=types.AudioTranscriptionConfig(),
        output_audio_transcription=types.AudioTranscriptionConfig(),
    )
    
  • Raw WebSockets: Place translationConfig inside generationConfig:
    {
      "setup": {
        "model": "models/gemini-3.5-live-translate-preview",
        "generationConfig": {
          "responseModalities": ["AUDIO"],
          "translationConfig": {
            "targetLanguageCode": "es",
            "echoTargetLanguage": true
          }
        }
      }
    }
    

Live Streaming Transcription (Gemini Live Transcribe)

The Live API supports real-time streaming speech-to-text over WebSockets with low-latency interim hypotheses, finalized transcripts, and Hybrid VAD. For full details, see the Live Transcription Guide and Colab Cookbook.

Model

  • gemini-3.5-transcribe-live

Modes

  • smart: cleans up filler words, resolves inline self-corrections, and structures formatting.
  • verbatim (default): exact word-for-word transcript.

Python

config = types.LiveConnectConfig(
    response_modalities=["TEXT"],
    input_audio_transcription=types.AudioTranscriptionConfig(),
)

async with client.aio.live.connect(model="gemini-3.5-transcribe-live", config=config) as session:
    # Stream audio
    await session.send_realtime_input(audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000"))
    # Hybrid VAD: notify turn end on client-detected silence for zero latency
    await session.send_realtime_input(audio_stream_end=True)

JavaScript

const session = await ai.live.connect({
  model: 'gemini-3.5-transcribe-live',
  config: {
    responseModalities: ['text'],
    inputAudioTranscription: { mode: 'smart' }
  },
  callbacks: {
    onmessage: (msg) => {
      if (msg.serverContent?.interimInputTranscription) {
        console.log('Interim:', msg.serverContent.interimInputTranscription.text);
      }
      if (msg.serverContent?.inputTranscription) {
        console.log('Final:', msg.serverContent.inputTranscription.text);
      }
    }
  }
});

session.sendRealtimeInput({ audio: { data: chunkBase64, mimeType: 'audio/pcm;rate=16000' } });
session.sendRealtimeInput({ audioStreamEnd: true }); // Hybrid VAD

Raw WebSockets

{
  "setup": {
    "model": "models/gemini-3.5-transcribe-live",
    "generationConfig": {
      "responseModalities": ["TEXT"],
      "speechConfig": {
        "voiceConfig": {}
      }
    },
    "inputAudioTranscription": {
      "mode": "smart"
    }
  }
}

Limitations

  • Response modality — Only TEXT or AUDIO per session, not both. Native audio models output audio (response_modalities=["AUDIO"]); enable output_audio_transcription if you need text transcripts.
  • Audio-only session — 15 min without compression
  • Audio+video session — 2 min without compression
  • Connection lifetime — ~10 min (use session resumption)
  • Context window — 128k input tokens / 64k output tokens
  • Code execution / URL context — Not supported

Upgrading & Migration

For step-by-step migration checklists and protocol deltas when upgrading from gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-*, or gemini-2.0-flash-live-001 to Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking, read references/migration.md.

Best Practices

  1. Use headphones when testing mic audio to prevent echo/self-interruption
  2. Enable context window compression for sessions longer than 15 minutes
  3. Implement session resumption to handle connection resets gracefully
  4. Use ephemeral tokens for client-side deployments — never expose API keys in browsers
  5. Use send_realtime_input for real-time user input (audio, video, text). Use send_client_content with explicit user/model roles to inject context turns mid-stream
  6. Send audioStreamEnd / audio_stream_end (Hybrid VAD) when the mic is paused or user finishes speaking
  7. Clear audio playback queues on interruption signals (interrupted: true)
  8. Process all parts in each server event — events can contain multiple content parts
  9. Monitor interaction_status (IN_PROGRESS vs IDLE) when using gemini-3.8-live-extended-thinking rather than relying on turn_complete alone

Documentation Lookup

When MCP is Installed (Preferred)

If the search_docs tool (from the Google MCP server) is available, use it as your only documentation source:

  1. Call search_docs with your query
  2. Read the returned documentation
  3. Trust MCP results as source of truth for API details — they are always up-to-date.

[!IMPORTANT] When MCP tools are present, never fetch URLs manually. MCP provides up-to-date, indexed documentation that is more accurate and token-efficient than URL fetching.

When MCP is NOT Installed (Fallback Only)

If no MCP documentation tools are available, fetch from the official docs index:

llms.txt URL: https://ai.google.dev/gemini-api/docs/llms.txt

This index contains links to all documentation pages in .md.txt format. Use web fetch tools to:

  1. Fetch llms.txt to discover available documentation pages
  2. Fetch specific pages (e.g., https://ai.google.dev/gemini-api/docs/live-session.md.txt)

Key Documentation Pages

[!IMPORTANT] Those are not all the documentation pages. Use the llms.txt index to discover available documentation pages

Supported Languages

The Live API supports 70 languages including: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Hindi, Arabic, Russian, and many more. Native audio models automatically detect and switch languages.

Plus de skills de google-gemini

agent-tui
google-gemini
Main Agents: Do NOT use this skill directly. If you need to test the TUI, invoke the `tui_tester` subagent. Drive terminal UI (TUI) applications…
gemini-api-cli
google-gemini
Guide pour utiliser l'outil en ligne de commande de l'API Gemini. À utiliser lorsque vous devez interagir avec l'API Gemini via la ligne de commande, gérer des agents ou générer du contenu multimédia (images,…)
behavioral-evals
google-gemini
Conseils pour créer, exécuter, corriger et promouvoir des évaluations comportementales. À utiliser lors de la vérification de la logique de décision d’un agent, du débogage d’échecs, du débogage d’invites…
gemini-omni-flash-api
google-gemini
Utilisez cette compétence pour le montage vidéo génératif, la conversion texte-vers-vidéo, la génération vidéo avec référence d'image, et les animations de transition de première image vers vidéo en utilisant le…
gemini-api-dev
google-gemini
Développez des applications avec les modèles Gemini de Google, prenant en charge le contenu multimodal, l'appel de fonctions et les sorties structurées en Python, JavaScript, Go et Java. Accédez aux modèles actuels Gemini 3 (Pro, Flash, Pro Image) avec un contexte de 1 million de tokens ; les modèles hérités Gemini 2.x et 1.5 sont obsolètes. Prend en charge la génération de texte, la compréhension d'images/audio/vidéo, l'appel de fonctions, la sortie JSON structurée, l'exécution de code, la mise en cache de contexte et les embeddings. Kits SDK officiels disponibles : google-genai (Python),...
gemini-interactions-api
google-gemini
Interface unifié pour les modèles et agents Gemini avec état côté serveur, streaming et orchestration d'outils. Prend en charge plusieurs modèles actuels (gemini-3-flash-preview, gemini-3-pro-preview, gemini-2.5-flash/pro) et l'agent Deep Research ; remplace automatiquement les ID de modèles obsolètes par des alternatives actuelles. Déchargez l'historique des conversations sur le serveur via previous_interaction_id pour des interactions multi-tours avec état sans gestion manuelle de l'historique. Orchestration d'outils intégrée incluant...
deliver
google-gemini
Publie une version condensée du briefing vers un webhook entrant Google Chat ou Slack, afin que le run quotidien se livre tout seul — passe silencieusement lorsqu'aucun webhook n'est…
fetch-news
google-gemini
Récupère les dernières actualités de Google News et Hacker News pour chaque sujet et genre correspondant aux intérêts du lecteur, avec déduplication par rapport à tous les éléments affichés lors des exécutions précédentes.