Google is adding three audio models to the Gemini API and AI Studio: Gemini 3.8 Live, an Extended Thinking variant and Gemini 3.5 Transcribe. They target fluid voice dialogue that can still perform tasks and understand visual input.

Quick answer

ModelMain use
Gemini 3.8 LiveLow-latency dialogue and actions.
Extended ThinkingComplex multi-step voice reasoning.
Gemini 3.5 TranscribeSpeech recognition across 85+ languages.

Voice as an agent interface

The system does more than transcribe and read a response. It manages interruptions, conversational pacing and tool calls during a session. Extended Thinking trades some latency for more complex work.

Specialized transcription

Gemini 3.5 Transcribe targets captions, minutes and indexing. Teams should measure names, accents, noise and language switching on their own recordings rather than trust one global score.

Safety and disclosure

Google says generated audio includes an imperceptible watermark. That can help identify synthetic media, but does not replace disclosure or secure handling of recordings and logs.

Production evaluation

Test end-to-end latency, interruption handling, tool failures and cost per conversation. Natural speech can make an error sound more authoritative than text, so sensitive actions still need explicit confirmation and traceability.