Gemini 3.5 Live Translate: smooth and natural voice translation
Google/DeepMind
Google DeepMind has released Gemini 3.5 Live Translate, a speech-to-speech translation model that detects 70+ languages and generates natural-sounding translated speech with preserved intonation. The model rolls out across Google products including the Gemini Live API, Google Meet, and Google Translate apps, with continuous streaming and noise robustness.
Google DeepMind announced Gemini 3.5 Live Translate, a new audio model for live speech-to-speech translation that automatically detects over 70 languages and generates smooth translated speech while preserving the speaker's intonation, pacing, and pitch. Unlike turn-by-turn systems, the model streams speech continuously, staying just a few seconds behind the speaker. The model is rolling out starting today across Google products: for developers via the Gemini Live API and Google AI Studio in public preview; for enterprises in private preview this month in Google Meet; and for consumers via Google Translate on Android and iOS. It processes multilingual inputs without manual configuration and is robust to noise. Developer platforms like Agora, Fishjam, LiveKit, Pipecat, and Vision Agents have integrated with the Gemini Live API to simplify building voice translation apps. Partner Grab is testing the model for multilingual communication between drivers and travelers. Companies like CJ ENM and LiveKit have given positive feedback on translation quality and low latency. In Google Meet, speech translation will expand from 5 languages to 70+ languages, enabling over 2000 language combinations. The Google Translate app introduces a 'listening mode' on Android that plays translations through the phone's earpiece. All generated audio is watermarked with SynthID to prevent misinformation.
- Сокращения
- API = Application Programming Interface — интерфейс программирования приложений
Source: Google DeepMind —
original
