Google Deep Mind독점

Gemini 3.1 Flash TTS (preview)

이 모델 비교

Google DeepMind에서 개발한 음성 합성 모델. 텍스트로부터 고품질 음성을 생성합니다.

파라미터

Undisclosed

컨텍스트

8K

라이선스

Proprietary

출시일

2026-04-16

벤치마크 성능

AA Intelligence Index

LMArena Elo

HLE

ARC-AGI-2

SWE-bench Verified

GPQA Diamond

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

일본어 처리 능력

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API 가격

이 모델의 API 가격 정보는 현재 공개되지 않았습니다

강점

    약점

      활용 사례

        심층 분석

        Arena Elo

        1211

        #2 overall on Artificial Analysis TTS leaderboard

        Input Price

        $1.00/1M tokens

        via Google AI Studio

        Output Price

        $20.00/1M tokens

        audio output tokens

        Languages Supported

        70+

        including major global languages

        Latency

        <200ms

        for short-form content generation

        Audio Tags

        200+

        for expressive control via inline natural language tags

        강점

        • Unmatched expressiveness with 200+ inline audio tags for precise voice control.
        • Low latency (<200ms) enabling real-time interactive applications.
        • Broad multilingual support covering 70+ languages for global use.

        약점

        • Preview status with potential instability, no SLA, and possible breaking changes.
        • Limited voice library with only 30 preset voices and no voice cloning support.
        • Significant quality degradation for long-form content over one minute, with 90% failure rate observed.

        경쟁사 비교

        ModelArenaSWEGPQAPrice
        ElevenLabs Flash#3N/AN/A$0.050/1K chars
        OpenAI TTS-1-HDBelow top 10N/AN/A$0.030/1K chars

        Gemini 3.1 Flash TTS, released by Google DeepMind in April 2026, is a preview text-to-speech model designed for high-fidelity, expressive speech synthesis. It stands out with granular control through 200+ inline audio tags, allowing developers to direct vocal style, pacing, and emotion directly in text prompts. The model supports 70+ languages, features native multi-speaker dialogue, and includes SynthID watermarking for AI-generated content detection.

        Positioned as a cost-effective alternative to premium TTS services, Gemini 3.1 Flash TTS excels in short-form applications like real-time narration, accessibility tools, and e-learning content. However, its preview nature and limitations in long-form stability and voice diversity position it primarily for prototyping and specific use cases where expressiveness and low cost outweigh production robustness needs.

        분석 생성일: 2026-07-17