Gemini 3.6 Flash is Google DeepMind's flagship efficient reasoning model, launched July 21, 2026. It runs a 1,000,000-token multimodal context (text, image, audio, video, PDF) and costs $1.50 per million input and $7.50 per million output tokens. Google reports it uses about 17% fewer output tokens than Gemini 3.5 Flash on the same tasks, while improving coding (SWE-Bench Pro 58.7%), agentic computer use (OSWorld-Verified 83.0%), and terminal work (Terminal-Bench 2.1 78%). It is the default Flash-tier workhorse for agentic coding, tool workflows, and long-document reasoning.
Parameters
Undisclosed
Context Window
1M
License
Proprietary
Release Date
2026-07-21
API Pricing
Input Price (per 1M tokens)
$1.5
Output Price (per 1M tokens)
$7.5
Billing Mode: standard
Strengths
- •Enhanced performance and capability over previous versions.
- •Improved handling of complex reasoning and instruction following.
- •Strong multimodal integration for text and other data types.
Weaknesses
- •Increased resource requirements compared to Flash-Lite variants.
- •Potential for increased latency in resource-constrained environments.
- •Complexity may introduce new edge cases in response generation.
Use Cases
Deep Analysis
Context Window
1M tokens
1,048,576 input / 65,536 output; text, image, audio, video, PDF
Input/Output Price
$1.50/$7.50 per 1M
Standard API rate; ~1.7x cheaper than Gemini 3.1 Pro
SWE-Bench Pro
58.7%
+3.6 pts over Gemini 3.5 Flash (55.1%)
Terminal-Bench 2.1
78.0%
Agentic terminal workflows
OSWorld-Verified
83.0%
Computer use; +4.6 pts over 3.5 Flash
Token Efficiency
17% fewer output
vs Gemini 3.5 Flash on same tasks
Strengths
- ・Runs a full 1M-token multimodal context (text, image, audio, video, PDF) at Flash-tier latency and cost
- ・Strong agentic and tool-use performance — OSWorld-Verified 83.0%, Terminal-Bench 2.1 78.0%, MLE-Bench 63.9%
- ・About 17% fewer output tokens than 3.5 Flash with better coding, agentic, and knowledge-work scores; ~1.7x cheaper than 3.1 Pro
Weaknesses
- ・Trails GPT-5.5 on Terminal-Bench (82.7% vs 78.0%) and Claude Opus 4.7 on OSWorld-Verified (82.8%) for heavy infrastructure work
- ・Reasoning ceilings below top reasoning models — HLE and ARC-AGI-style abstract tasks remain a gap versus Opus-class models
- ・Proprietary and Google-only; no open weights, and rate limits apply on high-volume agent loops without enterprise tiers
Competitor Comparison
| Model | SWE | Price |
|---|---|---|
| Gemini 3.6 Flash | 58.7% (SWE-Bench Pro) | $1.50/$7.50 per 1M |
| Gemini 3.5 Flash | 55.1% (SWE-Bench Pro) | $1.50/$9.00 per 1M |
| GPT-5.5 | 58.6% (SWE-Bench Pro) | ~$10/$30 per 1M |
| Claude Opus 4.7 | 64.3% (SWE-Bench Pro) | ~$15/$75 per 1M |
Gemini 3.6 Flash is Google DeepMind's flagship efficient reasoning model, launched July 21, 2026. It pairs a 1,000,000-token multimodal context (text, image, audio, video, and PDF) with Flash-tier latency and a standard API price of $1.50 per million input and $7.50 per million output tokens. Google reports it uses about 17% fewer output tokens than Gemini 3.5 Flash on the same tasks while improving coding (SWE-Bench Pro 58.7%), agentic computer use (OSWorld-Verified 83.0%), and terminal work (Terminal-Bench 2.1 78.0%). It is the default Flash-tier workhorse for agentic coding, tool workflows, and long-document reasoning.
Architecturally, 3.6 Flash continues Google's efficient-reasoning direction: strong tool calling, function calling, and computer use (screen understanding, UI actions) with fewer unwanted edit loops and hallucinations than 3.5 Flash. It leads the Flash family on MLE-Bench (63.9% vs 49.7%), CharXiv-R (89.4%), and MRCR v2 8-needle (54.0%), and holds a GDPval-AA agentic knowledge-work score around 1421. Against larger frontier models, it is roughly 1.7x cheaper than Gemini 3.1 Pro and far cheaper than Claude Opus 4.7 or GPT-5.5, while closing much of the coding and agentic gap.
The main trade-offs are ceiling and access. On Terminal-Bench 2.1 it trails GPT-5.5 (82.7%) by about 4 points, and on OSWorld-Verified it trails Claude Opus 4.7 (82.8%) by a hair. Abstract-reasoning and HLE-style tasks remain below Opus-class models. As a proprietary, Google-only model with no open weights, it is best deployed through Vertex AI or the Gemini API for secure, scalable production agent fleets.
Sources
Analysis generated: 2026-08-14