Google Deep MindProprietary

Gemini 3.5 Live Translate

Compare this model

The latest multimodal model from Google DeepMind, designed for integrated processing of text, images, and video.

Parameters

Undisclosed

Context Window

License

Proprietary

Release Date

2026-06-09

Japanese Language Capability

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API Pricing

Input Price (per 1M tokens)

$3.5

Output Price (per 1M tokens)

$

Billing Mode: standard

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        Languages Supported

        70+

        auto-detected; 2,000+ pair combinations in one session

        First-Audio Latency

        ~2.9 seconds

        median 2,947 ms (LiveLingo benchmark, p10–p90: 2,859–3,104 ms)

        Comprehension Score

        4.93 / 5

        #1 among tested systems (LiveLingo 2026 benchmark, 120 utterances)

        Effective Cost

        ~$0.037/min

        $3.50/1M input tokens, $21.00/1M output tokens

        Input Context

        128K tokens

        audio context window; 64K token output limit

        Base Architecture

        Gemini 3 Pro

        natively multimodal audio model, not a cascaded STT→MT→TTS pipeline

        Strengths

        • Native audio-to-audio architecture produces the most natural-sounding translated speech, preserving speaker intonation, pacing, and pitch
        • Broadest distribution of any live translation system—rolled into Google Translate, Google Meet, and the Gemini Live API simultaneously
        • Continuous streaming translation (~3s latency) enables real conversational flow instead of turn-by-turn pauses

        Weaknesses

        • Audio-only output with no streaming text mode, no per-speaker attribution, and no ability to edit or revise spoken output mid-utterance
        • Voice inconsistency documented by Google's own model card—voices can shift gender, get stuck, or bleed across speakers in multi-party scenarios
        • Structural blind spot on code-switched audio: when source speech switches into the target language, content silently disappears (~28% loss in benchmark testing)

        Competitor Comparison

        ModelArenaSWEGPQAPrice
        OpenAI gpt-realtime-translateN/AN/AN/A$0.06/min est.
        Google Cloud STT v2 + Translate v3N/AN/AN/Avaries per service
        Azure Speech TranslationN/AN/AN/Avaries per tier

        Gemini 3.5 Live Translate, released June 9, 2026, is Google DeepMind's specialized audio-to-audio translation model built on the Gemini 3 Pro architecture. Unlike traditional cascaded pipelines (speech-to-text → machine translation → text-to-speech), it processes audio natively—accepting 16kHz PCM chunks and outputting 24kHz translated speech in near real-time. The model auto-detects 70+ languages and preserves the speaker's vocal characteristics, producing translated audio that sounds substantially more natural than generic TTS reading a translation aloud. It ships simultaneously across Google Translate (consumer), Google Meet (enterprise), and the Gemini Live API (developer), giving Google an unmatched distribution advantage in the live translation space.

        The model represents a fundamental architectural shift in how translation is delivered. Rather than being a feature bolted onto existing products, it positions translation as an ambient capability—a layer that sits inside conversations, meetings, and apps. Google's integration with partners like Grab (10M+ monthly voice calls), Agora, LiveKit, and Pipecat signals that the long-term vision is embedded translation infrastructure, not a standalone translator app. At ~$0.037 per minute of translated conversation, the economics are viable for high-value business use cases while remaining accessible to developers through a free tier in Google AI Studio.

        However, the model card is refreshingly honest about limitations. Voice consistency degrades in multi-speaker sessions, language detection struggles with accents and similar language pairs, and the irreversible audio commitment means late-resolving syntax (common in Mandarin, Japanese) can produce factual inversions. The absence of streaming text output, per-speaker attribution, and mid-utterance revision makes it unsuitable for use cases requiring verbatim records or speaker disambiguation. This is a model optimized for fluid conversational translation—not transcription, not summarization, and not the kind of editable output that enterprise compliance workflows demand.

        Analysis generated: 2026-07-17