Google Adds Two Voice Models That Can See What You Show Them

Google released two live-dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, on Sept. 15, 2026, and made both available to developers the same day.
The company says that the pair is built for voice agents, meaning software that a person talks to instead of typing at, and that what changes is that the model no longer goes quiet while it works. Google's release says Gemini 3.8 Live runs tools and calls other software in the background, acknowledges the request out loud, and keeps the conversation going while those tasks finish.
The same model, Google says, takes visual input in near real time, so a user can show it something mid-sentence, and detects and switches between 97 supported languages during a conversation. The second model, 3.8 Live Extended Thinking, is the one Google puts forward for longer multi-step tasks: it reasons and speaks at once, the company says, narrating its progress while work runs in the background.
Every placement published with the release is Google's own account of where outside scoreboards put its models. Google reports that 3.8 Live Extended Thinking took the top overall spot on Artificial Analysis' Speech to Speech Quality Index, with a score of 82.6, and that 3.8 Live came second in the Speech Agent Arena.
Both models reached the Gemini API and Google AI Studio on the day of the announcement and are in private preview for enterprise customers, Google said. 3.8 Live is also going into Search Live; 3.8 Live Extended Thinking is going into Gemini Live, and into Docs, Gmail and Keep for Google AI subscribers.
Google watermarks all audio its AI products generate with SynthID, an inaudible marker the company says keeps machine-made audio detectable.
