Google Deep MindProprietary

Gemini 3.6 Flash

Compare this model

Gemini 3.6 Flash is Google DeepMind's flagship efficient reasoning model, launched July 21, 2026. It runs a 1,000,000-token multimodal context (text, image, audio, video, PDF) and costs $1.50 per million input and $7.50 per million output tokens. Google reports it uses about 17% fewer output tokens than Gemini 3.5 Flash on the same tasks, while improving coding (SWE-Bench Pro 58.7%), agentic computer use (OSWorld-Verified 83.0%), and terminal work (Terminal-Bench 2.1 78%). It is the default Flash-tier workhorse for agentic coding, tool workflows, and long-document reasoning.

Parameters

Undisclosed

Context Window

1M

License

Proprietary

Release Date

2026-07-21

API Pricing

Input Price (per 1M tokens)

$1.5

Output Price (per 1M tokens)

$7.5

Billing Mode: standard

Strengths

  • Enhanced performance and capability over previous versions.
  • Improved handling of complex reasoning and instruction following.
  • Strong multimodal integration for text and other data types.

Weaknesses

  • Increased resource requirements compared to Flash-Lite variants.
  • Potential for increased latency in resource-constrained environments.
  • Complexity may introduce new edge cases in response generation.

Use Cases

    Deep Analysis

    Context Window

    1M tokens

    1,048,576 input / 65,536 output; text, image, audio, video, PDF

    Input/Output Price

    $1.50/$7.50 per 1M

    Standard API rate; ~1.7x cheaper than Gemini 3.1 Pro

    SWE-Bench Pro

    58.7%

    +3.6 pts over Gemini 3.5 Flash (55.1%)

    Terminal-Bench 2.1

    78.0%

    Agentic terminal workflows

    OSWorld-Verified

    83.0%

    Computer use; +4.6 pts over 3.5 Flash

    Token Efficiency

    17% fewer output

    vs Gemini 3.5 Flash on same tasks

    Strengths

    • Runs a full 1M-token multimodal context (text, image, audio, video, PDF) at Flash-tier latency and cost
    • Strong agentic and tool-use performance — OSWorld-Verified 83.0%, Terminal-Bench 2.1 78.0%, MLE-Bench 63.9%
    • About 17% fewer output tokens than 3.5 Flash with better coding, agentic, and knowledge-work scores; ~1.7x cheaper than 3.1 Pro

    Weaknesses

    • Trails GPT-5.5 on Terminal-Bench (82.7% vs 78.0%) and Claude Opus 4.7 on OSWorld-Verified (82.8%) for heavy infrastructure work
    • Reasoning ceilings below top reasoning models — HLE and ARC-AGI-style abstract tasks remain a gap versus Opus-class models
    • Proprietary and Google-only; no open weights, and rate limits apply on high-volume agent loops without enterprise tiers

    Competitor Comparison

    ModelSWEPrice
    Gemini 3.6 Flash58.7% (SWE-Bench Pro)$1.50/$7.50 per 1M
    Gemini 3.5 Flash55.1% (SWE-Bench Pro)$1.50/$9.00 per 1M
    GPT-5.558.6% (SWE-Bench Pro)~$10/$30 per 1M
    Claude Opus 4.764.3% (SWE-Bench Pro)~$15/$75 per 1M

    Gemini 3.6 Flash is Google DeepMind's flagship efficient reasoning model, launched July 21, 2026. It pairs a 1,000,000-token multimodal context (text, image, audio, video, and PDF) with Flash-tier latency and a standard API price of $1.50 per million input and $7.50 per million output tokens. Google reports it uses about 17% fewer output tokens than Gemini 3.5 Flash on the same tasks while improving coding (SWE-Bench Pro 58.7%), agentic computer use (OSWorld-Verified 83.0%), and terminal work (Terminal-Bench 2.1 78.0%). It is the default Flash-tier workhorse for agentic coding, tool workflows, and long-document reasoning.

    Architecturally, 3.6 Flash continues Google's efficient-reasoning direction: strong tool calling, function calling, and computer use (screen understanding, UI actions) with fewer unwanted edit loops and hallucinations than 3.5 Flash. It leads the Flash family on MLE-Bench (63.9% vs 49.7%), CharXiv-R (89.4%), and MRCR v2 8-needle (54.0%), and holds a GDPval-AA agentic knowledge-work score around 1421. Against larger frontier models, it is roughly 1.7x cheaper than Gemini 3.1 Pro and far cheaper than Claude Opus 4.7 or GPT-5.5, while closing much of the coding and agentic gap.

    The main trade-offs are ceiling and access. On Terminal-Bench 2.1 it trails GPT-5.5 (82.7%) by about 4 points, and on OSWorld-Verified it trails Claude Opus 4.7 (82.8%) by a hair. Abstract-reasoning and HLE-style tasks remain below Opus-class models. As a proprietary, Google-only model with no open weights, it is best deployed through Vertex AI or the Gemini API for secure, scalable production agent fleets.

    Analysis generated: 2026-08-14