Google is adding three audio models to the Gemini API and AI Studio: Gemini 3.8 Live, an Extended Thinking variant and Gemini 3.5 Transcribe. They target fluid voice dialogue that can still perform tasks and understand visual input.
Quick answer
| Model | Main use |
|---|---|
| Gemini 3.8 Live | Low-latency dialogue and actions. |
| Extended Thinking | Complex multi-step voice reasoning. |
| Gemini 3.5 Transcribe | Speech recognition across 85+ languages. |
Voice as an agent interface
The system does more than transcribe and read a response. It manages interruptions, conversational pacing and tool calls during a session. Extended Thinking trades some latency for more complex work.
Specialized transcription
Gemini 3.5 Transcribe targets captions, minutes and indexing. Teams should measure names, accents, noise and language switching on their own recordings rather than trust one global score.
Safety and disclosure
Google says generated audio includes an imperceptible watermark. That can help identify synthetic media, but does not replace disclosure or secure handling of recordings and logs.
Production evaluation
Test end-to-end latency, interruption handling, tool failures and cost per conversation. Natural speech can make an error sound more authoritative than text, so sensitive actions still need explicit confirmation and traceability.




Join the discussion
Comments
Loading comments…