Google Deep Mind독점

Gemini 3.5 Live Translate

이 모델 비교

Google DeepMind가 개발한 최신 멀티모달 모델. 텍스트, 이미지, 영상을 통합적으로 처리하며 실시간 번역 기능을 제공합니다.

파라미터

Undisclosed

컨텍스트

라이선스

Proprietary

출시일

2026-06-09

벤치마크 성능

AA Intelligence Index

LMArena Elo

HLE

ARC-AGI-2

SWE-bench Verified

GPQA Diamond

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

일본어 처리 능력

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API 가격

입력 가격 (1M 토큰당)

$3.5

출력 가격 (1M 토큰당)

$

과금 모드: standard

강점

    약점

      활용 사례

        심층 분석

        Languages Supported

        70+

        auto-detected; 2,000+ pair combinations in one session

        First-Audio Latency

        ~2.9 seconds

        median 2,947 ms (LiveLingo benchmark, p10–p90: 2,859–3,104 ms)

        Comprehension Score

        4.93 / 5

        #1 among tested systems (LiveLingo 2026 benchmark, 120 utterances)

        Effective Cost

        ~$0.037/min

        $3.50/1M input tokens, $21.00/1M output tokens

        Input Context

        128K tokens

        audio context window; 64K token output limit

        Base Architecture

        Gemini 3 Pro

        natively multimodal audio model, not a cascaded STT→MT→TTS pipeline

        강점

        • Native audio-to-audio architecture produces the most natural-sounding translated speech, preserving speaker intonation, pacing, and pitch
        • Broadest distribution of any live translation system—rolled into Google Translate, Google Meet, and the Gemini Live API simultaneously
        • Continuous streaming translation (~3s latency) enables real conversational flow instead of turn-by-turn pauses

        약점

        • Audio-only output with no streaming text mode, no per-speaker attribution, and no ability to edit or revise spoken output mid-utterance
        • Voice inconsistency documented by Google's own model card—voices can shift gender, get stuck, or bleed across speakers in multi-party scenarios
        • Structural blind spot on code-switched audio: when source speech switches into the target language, content silently disappears (~28% loss in benchmark testing)

        경쟁사 비교

        ModelArenaSWEGPQAPrice
        OpenAI gpt-realtime-translateN/AN/AN/A$0.06/min est.
        Google Cloud STT v2 + Translate v3N/AN/AN/Avaries per service
        Azure Speech TranslationN/AN/AN/Avaries per tier

        Gemini 3.5 Live Translate, released June 9, 2026, is Google DeepMind's specialized audio-to-audio translation model built on the Gemini 3 Pro architecture. Unlike traditional cascaded pipelines (speech-to-text → machine translation → text-to-speech), it processes audio natively—accepting 16kHz PCM chunks and outputting 24kHz translated speech in near real-time. The model auto-detects 70+ languages and preserves the speaker's vocal characteristics, producing translated audio that sounds substantially more natural than generic TTS reading a translation aloud. It ships simultaneously across Google Translate (consumer), Google Meet (enterprise), and the Gemini Live API (developer), giving Google an unmatched distribution advantage in the live translation space.

        The model represents a fundamental architectural shift in how translation is delivered. Rather than being a feature bolted onto existing products, it positions translation as an ambient capability—a layer that sits inside conversations, meetings, and apps. Google's integration with partners like Grab (10M+ monthly voice calls), Agora, LiveKit, and Pipecat signals that the long-term vision is embedded translation infrastructure, not a standalone translator app. At ~$0.037 per minute of translated conversation, the economics are viable for high-value business use cases while remaining accessible to developers through a free tier in Google AI Studio.

        However, the model card is refreshingly honest about limitations. Voice consistency degrades in multi-speaker sessions, language detection struggles with accents and similar language pairs, and the irreversible audio commitment means late-resolving syntax (common in Mandarin, Japanese) can produce factual inversions. The absence of streaming text output, per-speaker attribution, and mid-utterance revision makes it unsuitable for use cases requiring verbatim records or speaker disambiguation. This is a model optimized for fluid conversational translation—not transcription, not summarization, and not the kind of editable output that enterprise compliance workflows demand.

        분석 생성일: 2026-07-17