Google Launches Gemini 3.8 Live and Live Extended Thinking for Real-Time Voice
What launched
FAQ
What is Gemini 3.8 Live?
It is a speech-to-speech model from Google that converts audio directly to audio without a text intermediary, reducing latency and allowing natural conversation with mid-speech interruption.
How does Gemini 3.8 Live compare to OpenAI's GPT-Live?
Both are bidirectional voice models. Gemini 3.8 Live is exposed through a WebSocket endpoint in the Google AI API with an Extended Thinking variant for deeper reasoning, while GPT-Live runs on OpenAI interfaces. Practical differences show up in Arabic dialect quality, pricing, and Google Cloud integration.
Should MENA enterprises adopt Gemini 3.8 Live now?
It suits early pilots and multilingual contact centers, but test multiple Arabic dialects before production and review data governance and local compliance requirements per country.
What does Extended Thinking mean in a voice model?
It gives the model more reasoning time before responding, useful for complex multi-step questions, though it may increase response latency versus the real-time mode.
Source: Simon Willison (LLM & tools)
AI-assisted content, human-reviewed.