Google Deep MindProprietary

Gemini 3.5 Flash

Compare this model

The latest multimodal model developed by Google DeepMind. It performs integrated processing of text, images, and video.

Parameters

Undisclosed

Context Window

1M

License

Proprietary

Release Date

2026-06-20

Japanese Language Capability

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API Pricing

Input Price (per 1M tokens)

$1.5

Output Price (per 1M tokens)

$

Billing Mode: standard

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        Arena Elo

        1479

        #8 overall in Text Arena

        SWE-Bench Verified

        80.8%

        per BenchLM aggregator; Google reports 55.1% on SWE-Bench Pro

        Input Price

        $1.50/1M

        3x higher than Gemini 3 Flash; ~40% lower than Gemini 3.1 Pro

        Output Speed

        ~284 tok/s

        ~4x faster than frontier peers

        Context Window

        1M tokens

        Weak recall at full 1M (26.6% on MRCR)

        Key Agentic Benchmark

        83.6%

        Field-leading MCP Atlas score

        Strengths

        • Best-in-class agentic tool use and multi-step workflow performance
        • Exceptional speed (~4x faster than frontier models) at a competitive price
        • Strong native multimodal capabilities including video and chart reasoning

        Weaknesses

        • 3x price increase over previous Flash generation breaks budget assumptions
        • Poor long-context recall (26.6% at 1M tokens) despite the 1M window claim
        • Lacks computer-use capability and trails on pure reasoning benchmarks vs. Pro

        Competitor Comparison

        ModelArenaSWEGPQAPrice
        GPT-5.5 (OpenAI)~176958.6% (Pro)~92%$5/$30 per 1M
        Claude Opus 4.7 (Anthropic)~175364.3% (Pro)~91%~$15/$75 per 1M
        Gemini 3.1 Pro (Google)~131454.2% (Pro)94.3%$2/$12 per 1M

        Gemini 3.5 Flash represents a strategic repositioning of Google's 'Flash' tier from a budget model to a frontier-class agentic specialist. Launched at Google I/O 2026, it delivers state-of-the-art performance in multi-step tool use (MCP Atlas: 83.6%) and agentic coding (Terminal-Bench 2.1: 76.2%), often beating the previous Gemini 3.1 Pro while running approximately 4x faster. This makes it the default production backbone for coding agents, computer automation, and multimodal workflows within the Google ecosystem.

        However, this upgrade comes with a significant 3x price increase over its predecessor, moving Flash into a mid-tier pricing bracket ($1.50/$9 per million tokens). While still 40% cheaper than Gemini 3.1 Pro, it is no longer the budget option. The model's key innovation is its thought preservation feature for multi-turn conversations and adjustable reasoning levels, but it has notable weaknesses: the 1M context window has poor recall at full length, it lacks computer-use support, and it trails Pro models on pure reasoning benchmarks like Humanity's Last Exam (40.2% vs. 44.4%).

        In summary, Gemini 3.5 Flash is the premier choice for high-volume, speed-sensitive agentic applications on the Google stack, but developers must recalibrate expectations and budgets away from the old Flash tier's economics and evaluate whether Pro-tier models are still needed for the hardest reasoning tasks.

        Analysis generated: 2026-07-17