Google's native audio-to-audio models of the 3.8 generation, released 15 September 2026 as two API models: Gemini 3.8 Live (gemini-3.8-live) and Gemini 3.8 Live Extended Thinking (gemini-3.8-live-extended-thinking). Both speak 97 languages with mid-conversation switching, make asynchronous tool calls while continuing to talk, and watermark their audio with SynthID. Google reports 68.6% on τ-Voice and 97.7% on Big Bench Audio. Pricing is $3.00 per million audio input tokens and $12.00 per million audio output, or $0.75 and $4.50 for text. No parameter count or context length is published. On the Artificial Analysis Speech-to-Speech Index the Extended Thinking model scores 82.6, first place, and Live 76.0, ahead of OpenAI's GPT-Live-1 at 81.5; neither has an Intelligence Index page.

Model Details

License Proprietary (API)

Variants

Name Parameters Notes
Gemini 3.8 Live Extended Thinking AA Speech-to-Speech Index 82.6 (#1, Sep 2026).
Gemini 3.8 Live AA Speech-to-Speech Index 76.0.
audiospeechmultimodalproprietaryagentic

Related