Google Deep MindProprietary

Gemini 3.1 Pro Preview

Compare this model

A high-performance multimodal model developed by Google DeepMind. It performs integrated processing of text, images, and video.

Parameters

Undisclosed

Context Window

1M

License

Proprietary

Release Date

2026-02-20

Benchmark Performance

AA Intelligence Index

LMArena Elo

HLE

ARC-AGI-2

SWE-bench Verified

GPQA Diamond

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

Japanese Language Capability

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API Pricing

Input Price (per 1M tokens)

$2

Output Price (per 1M tokens)

$

Billing Mode: standard

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        GPQA Diamond

        94.3%

        #1 proprietary model (no tools)

        ARC-AGI-2

        77.1%

        ARC Prize verified; 2.5x jump from Gemini 3 Pro

        SWE-Bench Verified

        80.6%

        Single attempt; within ~0.2pp of Claude Opus 4.6

        LMArena Elo

        1481–1501

        Top 3 overall per multiple trackers

        Context Window

        1M tokens (2M on Vertex)

        Largest production context in market

        Input/Output Price

        $2.00 / $12.00 per 1M

        Standard tier (≤200K input); ~40% cheaper than Opus output

        Strengths

        • Highest proprietary GPQA Diamond score (94.3%) with class-leading scientific and multi-step reasoning
        • True 1M-token context with strong long-context retrieval (MRCR v2 128K: 84.9%), eliminating complex RAG pipelines
        • Best-in-class multimodal understanding spanning text, images, audio, and video up to 3 hours

        Weaknesses

        • Still in Preview status with no production SLA, potential API changes, and no confirmed GA date
        • Long-context pricing doubles above 200K input tokens ($4/$18), making full-window use expensive
        • High time-to-first-token (~28–30s) due to thinking phase limits snappy interactive use cases

        Competitor Comparison

        ModelArenaSWEGPQAPrice
        Claude Opus 4.5146276.3%89.1%$15/$75
        GPT-5.2~145280.0%92.4%$10/$30
        Gemini 3 Pro147076.2%91.9%$2/$12

        Gemini 3.1 Pro Preview, released February 19, 2026 by Google DeepMind, represents the largest single-version reasoning leap in the Gemini family's history. On ARC-AGI-2—a benchmark designed to resist memorization—the model jumped from 31.1% (Gemini 3 Pro) to a verified 77.1%, more than doubling performance in roughly three months. Paired with a record 94.3% on GPQA Diamond (graduate-level science reasoning) and a frontier-class 80.6% on SWE-Bench Verified, the model establishes Google as a clear top-tier contender across reasoning, coding, and multimodal tasks.

        The model's defining capabilities extend beyond raw benchmark numbers. Its native 1M-token context window (expandable to 2M on Vertex AI) is the largest in production as of mid-2026, with verified long-context retrieval accuracy of 84.9% at 128K tokens. Video understanding scores 87.2% on VideoMME—an 8-point gap over Claude Opus 4.5—making it the strongest option for video-heavy workflows. The model also benefits from deep integration with Google's ecosystem: Search grounding for live data, Workspace and BigQuery connectors, and enterprise-grade governance on Vertex AI.

        The primary caveats are the Preview label (no production SLA commitment, potential API surface changes before GA) and a tiered pricing structure that doubles costs for inputs exceeding 200K tokens. For teams already on Google Cloud or requiring frontier reasoning with multimodal and long-context support, Gemini 3.1 Pro Preview offers the strongest price-to-capability ratio in its tier. However, Claude Opus leads on pure coding autonomy and DeepSeek R2 on pure math reasoning, making workload-specific benchmarking essential before migration.

        Analysis generated: 2026-07-17