Overview
Gemini 3.1 Pro Preview, released February 19, 2026 by Google DeepMind, represents the largest single-version reasoning leap in the Gemini family's history. On ARC-AGI-2—a benchmark designed to resist memorization—the model jumped from 31.1% (Gemini 3 Pro) to a verified 77.1%, more than doubling performance in roughly three months. Paired with a record 94.3% on GPQA Diamond (graduate-level science reasoning) and a frontier-class 80.6% on SWE-Bench Verified, the model establishes Google as a clear top-tier contender across reasoning, coding, and multimodal tasks.
The model's defining capabilities extend beyond raw benchmark numbers. Its native 1M-token context window (expandable to 2M on Vertex AI) is the largest in production as of mid-2026, with verified long-context retrieval accuracy of 84.9% at 128K tokens. Video understanding scores 87.2% on VideoMME—an 8-point gap over Claude Opus 4.5—making it the strongest option for video-heavy workflows. The model also benefits from deep integration with Google's ecosystem: Search grounding for live data, Workspace and BigQuery connectors, and enterprise-grade governance on Vertex AI.
The primary caveats are the Preview label (no production SLA commitment, potential API surface changes before GA) and a tiered pricing structure that doubles costs for inputs exceeding 200K tokens. For teams already on Google Cloud or requiring frontier reasoning with multimodal and long-context support, Gemini 3.1 Pro Preview offers the strongest price-to-capability ratio in its tier. However, Claude Opus leads on pure coding autonomy and DeepSeek R2 on pure math reasoning, making workload-specific benchmarking essential before migration.
Benchmarks & Performance
Gemini 3.1 Pro Preview delivers state-of-the-art or near-state-of-the-art results across nearly every major benchmark category. The following table compiles official scores from Google DeepMind's model card and technical report, supplemented by independent evaluations from Artificial Analysis, evals.report, and ARC Prize verification.
| Benchmark | Category | Gemini 3.1 Pro | Gemini 3 Pro | Claude Opus 4.6 (Max) | GPT-5.2 (xhigh) | Notes |
|---|---|---|---|---|---|---|
| GPQA Diamond | Scientific reasoning | 94.3% | 91.9% | 91.3% | 92.4% | No tools; #1 proprietary |
| ARC-AGI-2 | Abstract reasoning | 77.1% | 31.1% | 68.8% | 52.9% | ARC Prize verified |
| Humanity's Last Exam (no tools) | Academic reasoning | 44.4% | 37.5% | 40.0% | 34.5% | Text + multimodal |
| Humanity's Last Exam (search+code) | Academic reasoning | 51.4% | 45.8% | 53.1% | 45.5% | With tools |
| SWE-Bench Verified | Agentic coding | 80.6% | 76.2% | 80.8% | 80.0% | Single attempt |
| SWE-Bench Pro (Public) | Diverse coding | 54.2% | 43.3% | — | 55.6% | Single attempt |
| LiveCodeBench Pro | Competitive coding | 2887 Elo | 2439 Elo | — | 2393 Elo | Codeforces/ICPC/IOI |
| Terminal-Bench 2.0 | Agentic terminal | 68.5% | 56.9% | 65.4% | 54.0% | Terminus-2 harness |
| AIME 2025/2026 | Math reasoning | 91.2–98.3% | 86.0% | — | — | Deep Think mode |
| τ2-bench Retail | Tool use | 90.8% | 85.3% | 91.9% | 82.0% | Agentic benchmark |
| τ2-bench Telecom | Tool use | 99.3% | 98.0% | 99.3% | 98.7% | Near-ceiling |
| MCP Atlas | Multi-step workflows | 69.2% | 54.1% | 59.5% | 60.6% | MCP-based |
| BrowseComp | Agentic search | 85.9% | 59.2% | 84.0% | 65.8% | Search + Python + Browse |
| MMMU-Pro | Multimodal reasoning | 80.5% | 81.0% | 73.9% | 79.5% | No tools |
| MMMLU | Multilingual Q&A | 92.6% | 91.8% | 91.1% | 89.6% | |
| MRCR v2 (128K) | Long context | 84.9% | 77.0% | 84.0% | 83.8% | 8-needle retrieval |
| MRCR v2 (1M) | Long context | 26.3% | 26.3% | N/A | N/A | Only Gemini supports 1M |
| SciCode | Scientific coding | 59% | 56% | 52% | 52% | |
| APEX-Agents | Long-horizon tasks | 33.5% | 18.4% | 29.8% | 23.0% | |
| VideoMME | Video understanding | 87.2% | 84.8% | ~79.2% | ~81.3% | 3-hour clips supported |
**Composite Indices:**
- Artificial Analysis Intelligence Index: 46–57 (depending on version; well above average)
- LMArena Elo: 1481–1501 (top 3 globally)
- Epoch Capabilities Index: 156.6
**Speed & Latency (via Artificial Analysis):**
- Output speed: ~116–143 tokens/second (above average for reasoning models)
- Time-to-first-token: ~28–30s (high; dominated by thinking phase)
- End-to-end response time for 500 tokens: notable latency due to TTFT
The model's strongest differentiators versus Gemini 3 Pro are in ARC-AGI-2 (+46pp), Terminal-Bench 2.0 (+11.6pp), LiveCodeBench Pro (+448 Elo), and MCP Atlas (+15.1pp)—all reflecting substantial gains in reasoning depth and agentic tool use.
Detailed Comparison
### vs Claude Opus 4.5
| Dimension | Gemini 3.1 Pro Preview | Claude Opus 4.5 |
|---|---|---|
| GPQA Diamond | 94.3% | 89.1% |
| SWE-Bench Verified | 80.6% | 76.3% (WhatLLM) / 80.8% (Google card) |
| ARC-AGI-2 | 77.1% | 68.8% |
| VideoMME | 87.2% | 79.2% |
| Context window | 1M tokens | 200K tokens |
| Input price | $2.00/1M | $15.00/1M |
| Output price | $12.00/1M | $75.00/1M |
Gemini 3.1 Pro dominates on reasoning depth, video understanding, context window, and cost—roughly 6x cheaper on output. Claude Opus 4.5 retains advantages in creative writing quality, instruction-following style, and some coding edge on multi-file refactor tasks. For teams running autonomous coding agents without budget constraints, Claude's SWE-Bench edge (~5pp per WhatLLM) translates to roughly 5 more resolved GitHub issues per 100. For everyone else, Gemini's pricing and context advantages are decisive.
### vs GPT-5.2
| Dimension | Gemini 3.1 Pro Preview | GPT-5.2 |
|---|---|---|
| GPQA Diamond | 94.3% | 92.4% |
| ARC-AGI-2 | 77.1% | 52.9% |
| SWE-Bench Verified | 80.6% | 80.0% |
| Terminal-Bench 2.0 | 68.5% | 54.0% |
| Context window | 1M tokens | 1M tokens |
| Input price | $2.00/1M | $10.00/1M |
| Output price | $12.00/1M | $30.00/1M |
Gemini 3.1 Pro outperforms GPT-5.2 on virtually every benchmark while costing 2.5x less on input and output. The ARC-AGI-2 gap (77.1% vs 52.9%) is particularly stark. GPT-5.2 retains advantages in tool-call reliability (per swfte.com report: 97.4% vs 91.6% success rate) and some enterprise workflows. OpenAI's ecosystem integrations (ChatGPT, DALL-E, Codex) remain a draw for teams already embedded in that stack.
### vs DeepSeek R2
| Dimension | Gemini 3.1 Pro Preview | DeepSeek R2 |
|---|---|---|
| AIME 2025 | 91.2% | 93.8% |
| GPQA Diamond | ~87.8% (WhatLLM) | 82.4% |
| SWE-Bench Verified | 71.4% (WhatLLM) | 62.1% |
| Video understanding | 87.2% (VideoMME) | Not supported |
| Context window | 1M tokens | 128K tokens |
| Input price | $2.00/1M | $0.55/1M |
| Output price | $12.00/1M | $2.19/1M |
DeepSeek R2 is the pure math champion (AIME 93.8%) and costs less than half on input. However, it is text-only (no image, audio, or video support), has a much smaller context window, and trails on agentic coding and tool use. For pure quantitative research on a budget, DeepSeek R2 wins. For multimodal, long-context, or production agentic workflows, Gemini 3.1 Pro is the superior choice.
Community Feedback
Developer and researcher reception has been broadly positive, focused on the verified ARC-AGI-2 leap and the pricing-to-capability ratio. Key themes from the community:
**Enthusiasm:** The ARC-AGI-2 jump from 31.1% to 77.1%—verified by ARC Prize—generated significant excitement as one of the largest single-version benchmark improvements in frontier model history. The GPQA Diamond score (94.3%) is widely cited as the new ceiling for scientific reasoning among proprietary models. Developers on the Gemini API praised the clean SDK experience (Python, TypeScript, Go, Java, Dart) and reliable function calling/structured output. The 1M-token context window is particularly popular among legal-tech, research-synthesis, and codebase-analysis teams who previously had to chunk documents for Claude or GPT.
**Skepticism:** The Preview label draws consistent criticism. Multiple reviewers (benchr, TopReviewed.ai's Skeptic persona, WhatLLM) note that building production systems on a preview checkpoint without SLA guarantees is a calculated risk. The 200K-token pricing cliff—where the entire request reprices to $4/$18 once input exceeds 200K—is flagged as a common cost surprise, especially for teams migrating from flat-priced models. The high TTFT (~30s) in Deep Think mode is noted as problematic for interactive applications.
**Adoption patterns:** The model has seen rapid adoption in Google Cloud-native enterprises already using Vertex AI, particularly for long-document analysis, multimodal pipelines, and scientific research assistance. The Workspace and BigQuery grounding capabilities are cited as a differentiator no competitor matches. Developer adoption on the API side is growing but tempered by rate limits on the Preview tier, which require cumulative spend to unlock higher thresholds. Teams on AWS or Azure are slower to adopt due to ecosystem lock-in concerns.
**Competitive framing:** The community generally positions Gemini 3.1 Pro as the best general-purpose frontier model ("best across the broadest range of tasks at the lowest cost" per WhatLLM), while acknowledging Claude Opus leads in coding autonomy and DeepSeek R2 in pure math. The imminent arrival of Gemini 3.5 Pro (announced at Google I/O, May 2026) has caused some procurement hesitancy, as 3.5 Flash already beats 3.1 Pro on several agentic benchmarks.
Use Cases
### 1. Long-Document Analysis and Research Synthesis
Gemini 3.1 Pro's 1M-token context window (2M on Vertex) makes it uniquely suited for processing entire codebases, legal contract sets, academic paper collections, or regulatory documents in a single prompt. At MRCR v2 84.9% accuracy at 128K tokens, it retrieves buried facts reliably. Example: a legal team ingests a 600K-token set of contracts and asks the model to identify all indemnification clauses with cross-references—something that would require complex RAG pipelines on Claude (200K limit) or GPT-5.2. The key cost caveat: inputs above 200K tokens trigger the $4/$18 long-context tier, so aggressive context trimming and retrieval engineering still pay off.
### 2. Scientific Research and Graduate-Level Q&A
With the highest GPQA Diamond score (94.3%) of any proprietary model, Gemini 3.1 Pro excels at multi-step scientific reasoning: composing 3-4 derivation steps in physics, biology, or chemistry while maintaining accuracy through the full chain. Example: a pharmaceutical researcher queries the model on drug interaction mechanisms requiring cross-referencing molecular pathways—it outperforms Claude Opus by ~5pp and GPT-5.2 by ~2pp on sustained scientific inference. Combined with Search grounding for literature updates, it serves as a powerful research assistant.
### 3. Video and Multimodal Pipeline Processing
At 87.2% on VideoMME, Gemini 3.1 Pro holds an 8-point advantage over Claude Opus on video understanding, supporting clips up to 3 hours. Use cases include: podcast summarization with temporal context, meeting transcription and action-item extraction, video content search and indexing, and video-to-code workflows (converting visual UI designs into React components). No other frontier model matches this capability at this price point.
### 4. Cost-Sensitive Frontend Development and UI Prototyping
Gemini 3.1 Pro ranks #1 on WebDev Arena (Elo ~1,443) for front-end code generation: transforming Figma mockups or screenshot references into React components, writing responsive CSS, and building interactive data visualizations. At $2/$12 per 1M tokens (vs Claude Opus at $15/$75), teams can run 6x more iterations for the same budget. Example: a product team provides a design mockup and the model generates a full interactive prototype with animations, 3D simulations, and live telemetry dashboards—as demonstrated in Google's own showcase.
**When to choose Gemini 3.1 Pro over alternatives:**
- Over Claude Opus: when cost, context window, video understanding, or scientific reasoning depth are primary constraints
- Over GPT-5.2: when multimodal capabilities, long context, or pricing matter more than tool-call reliability
- Over DeepSeek R2: when you need multimodal inputs, agentic workflows, or production-grade coding beyond pure math
- Over Gemini 3.5 Flash: when pure reasoning depth (HLE, GPQA) or long-context recall is more important than speed and agentic tool benchmarks
Latest News
**February 19, 2026:** Gemini 3.1 Pro Preview launched, featuring the largest reasoning leap in Gemini history. Key improvements include ARC-AGI-2 jumping from 31.1% to 77.1% (verified by ARC Prize), GPQA Diamond reaching 94.3%, and SWE-Bench Verified at 80.6%. Available via Google AI Studio, Gemini API, and Vertex AI in preview.
**February 2026:** A separate `gemini-3.1-pro-preview-customtools` endpoint was announced for developers building with bash and custom tool combinations, optimized for agentic workflows that use tools like `view_file` or `search_code`.
**May 19, 2026 (Google I/O):** Google launched Gemini 3.5 Flash to general availability and announced that Gemini 3.5 Pro is in testing with expected release in June 2026. Google positioned 3.5 Flash as already beating 3.1 Pro on coding, agentic, and multimodal benchmarks—signaling that 3.1 Pro's time as Google's flagship may be measured in weeks rather than quarters.
**May 2026 (ongoing):** The 2M-token extended context window is rolling out on Vertex AI for enterprise customers. The model remains in Preview status as of late May 2026 with no confirmed GA date. Google has been reportedly flexible on Preview-tier pricing commitments for enterprise customers willing to write case studies at GA.
**Pricing notes:** Standard rates are $2 input / $12 output per 1M tokens (≤200K input), jumping to $4/$18 above 200K. Batch API available at ~50% discount ($1/$6 under 200K). Cache hit pricing at $0.20/1M tokens (90% discount). Search grounding offers 5,000 free prompts/month across the Gemini 3 family, then $14 per 1,000 queries.