Gemini 3.5 Live Translate
Gemini 3.5 Live Translate is a near real-time speech-to-speech translation model released by Google on June 9, 2026. Its core function is to automatically detect the language spoken by the user and continuously generate translated audio in over 70 languages and more than 2,000 language pairs. While typical speech translation involves multiple steps that are visible to the user, such as converting speech to text, translating, and then synthesizing it back into speech, this model translates before the speaker has completely finished speaking, similar to an interpreter following along and immediately conveying the sentence. It takes into account the input speaker's intonation, speaking speed, and pitch.
Gemini 3.5 Live Translate is a near real-time speech-to-speech translation model released by Google on June 9, 2026. Its core function is to automatically detect the language spoken by the user and continuously generate translated audio in over 70 languages and 2,000+ language pairs. While typical speech translation reveals multiple steps to the user, such as converting speech to text, translating, and then synthesizing it back into speech, this model proceeds with translation even before the speaker has finished their utterance, similar to an interpreter following along and immediately conveying the sentence. It is designed to preserve the speaker's intonation, speaking speed, and pitch, focusing on maintaining not only the meaning of the translated text but also the manner of delivery.
Existing automatic translation can disrupt natural conversations due to delays in waiting for the sentence to end, loss of rhythm during step-by-step processing, and reduced recognition due to ambient noise. In situations where the tone itself is important, such as the emphasis of a presenter or the anxious voice of a patient, simply conveying accurate text can result in some information being lost. Continuous streaming translation differentiates itself by reducing these disruptions and maintaining the flow of conversation with output that reflects the speaking characteristics of the original speaker. Publicly available information indicating support for noisy environments also demonstrates its potential for use in uncontrolled settings such as on-site interviews, international events, and conversations while traveling.
Life science researchers can utilize it to provide real-time translation of discussions between researchers speaking different languages during multinational collaborative research meetings. For example, by sharing microscope observation results or clinical trial operational issues on a screen and connecting Live Translate to the audio channel, a workflow can be established to deliver automatically detected speech from one of the 70+ supported languages to the other researcher's language. The feature of preserving the speaker's speed and intonation can be useful in conveying the intent of utterances, such as questions, rebuttals, and warning signals, more naturally than with typical monotonous speech synthesis. However, further confirmation through official documentation is needed regarding the integration method with meeting platforms, the scope of API provision, supported regions, and data processing policies.
It can also be considered as an auxiliary translation layer for international patient interviews or multilingual field studies. Researchers can preserve the original audio separately and use Live Translate as a real-time channel to assist the flow of conversation, while using professional translation or verified transcripts as reference materials. The 2,000+ language pairs offer the possibility of operation without requiring an intermediate common language for each participant, but the accuracy of medical terminology, speaker separation, and the performance of number and drug name processing could not be confirmed based solely on the provided information. Therefore, it should be evaluated as a research aid that includes human review and comparison with the original text, rather than being used alone for clinical decision-making or obtaining consent.
💻 System Requirements
To be confirmed
To be confirmed
⚡ Installation
4-1. Quick Start
Not written because the official installation command and SDK package name were not found in the provided Discovery information.
4-2. Detailed Installation
After verifying the official API documentation, authentication methods, supported platforms, and usage permissions, additional details should be added. Unverified package names or arbitrary installation commands have not been included.
FAQ
What is Gemini 3.5 Live Translate?
Gemini 3.5 Live Translate is a near real-time speech-to-speech translation model released by Google on June 9, 2026. Its core function is to automatically detect the language spoken by the user and continuously generate translated audio in over 70 languages and 2,000+ language pairs. While typical speech translation reveals multiple steps to the user, such as converting speech to text, translating, and then synthesizing it back into speech, this model proceeds with translation even before the speaker has finished their utterance, similar to an interpreter following along and immediately conveying the sentence. It is designed to preserve the speaker's intonation, speaking speed, and pitch, focusing on maintaining not only the meaning of the translated text but also the manner of delivery. Existing automatic translation can disrupt natural conversations due to delays in waiting for the sentence to end, loss of rhythm during step-by-step processing, and reduced recognition due to ambient noise. In situations where the tone itself is important, such as the emphasis of a presenter or the anxious voice of a patient, simply conveying accurate text can result in some information being lost. Continuous streaming translation differentiates itself by reducing these disruptions and maintaining the flow of conversation with output that reflects the speaking characteristics of the original speaker. Publicly available information indicating support for noisy environments also demonstrates its potential for use in uncontrolled settings such as on-site interviews, international events, and conversations while traveling. Life science researchers can utilize it to provide real-time translation of discussions between researchers speaking different languages during multinational collaborative research meetings. For example, by sharing microscope observation results or clinical trial operational issues on a screen and connecting Live Translate to the audio channel, a workflow can be established to deliver automatically detected speech from one of the 70+ supported languages to the other researcher's language. The feature of preserving the speaker's speed and intonation can be useful in conveying the intent of utterances, such as questions, rebuttals, and warning signals, more naturally than with typical monotonous speech synthesis. However, further confirmation through official documentation is needed regarding the integration method with meeting platforms, the scope of API provision, supported regions, and data processing policies. It can also be considered as an auxiliary translation layer for international patient interviews or multilingual field studies. Researchers can preserve the original audio separately and use Live Translate as a real-time channel to assist the flow of conversation, while using professional translation or verified transcripts as reference materials. The 2,000+ language pairs offer the possibility of operation without requiring an intermediate common language for each participant, but the accuracy of medical terminology, speaker separation, and the performance of number and drug name processing could not be confirmed based solely on the provided information. Therefore, it should be evaluated as a research aid that includes human review and comparison with the original text, rather than being used alone for clinical decision-making or obtaining consent.
When should I use Gemini 3.5 Live Translate?
Gemini 3.5 Live Translate is a near real-time speech-to-speech translation model released by Google on June 9, 2026. Its core function is to automatically detect the language spoken by the user and continuously generate translated audio in over 70 languages and more than 2,000 language pairs. While typical speech translation involves multiple steps that are visible to the user, such as converting speech to text, translating, and then synthesizing it back into speech, this model translates before the speaker has completely finished speaking, similar to an interpreter following along and immediately conveying the sentence. It takes into account the input speaker's intonation, speaking speed, and pitch.
📝 Update Notes
No update notes yet.
🧪 Related Code of Life
No related Code of Life posts yet.