The latest multimodal model developed by Google DeepMind. It performs integrated processing of text, images, and video.
Parameters
Undisclosed
Context Window
1M
License
Proprietary
Release Date
2026-06-20
Japanese Language Capability
Multilingual model with strong Japanese language processing capabilities.
API Pricing
Input Price (per 1M tokens)
$1.5
Output Price (per 1M tokens)
$
Billing Mode: standard
Strengths
Weaknesses
Use Cases
Deep Analysis
Arena Elo
1479
#8 overall in Text Arena
SWE-Bench Verified
80.8%
per BenchLM aggregator; Google reports 55.1% on SWE-Bench Pro
Input Price
$1.50/1M
3x higher than Gemini 3 Flash; ~40% lower than Gemini 3.1 Pro
Output Speed
~284 tok/s
~4x faster than frontier peers
Context Window
1M tokens
Weak recall at full 1M (26.6% on MRCR)
Key Agentic Benchmark
83.6%
Field-leading MCP Atlas score
Strengths
- ・Best-in-class agentic tool use and multi-step workflow performance
- ・Exceptional speed (~4x faster than frontier models) at a competitive price
- ・Strong native multimodal capabilities including video and chart reasoning
Weaknesses
- ・3x price increase over previous Flash generation breaks budget assumptions
- ・Poor long-context recall (26.6% at 1M tokens) despite the 1M window claim
- ・Lacks computer-use capability and trails on pure reasoning benchmarks vs. Pro
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| GPT-5.5 (OpenAI) | ~1769 | 58.6% (Pro) | ~92% | $5/$30 per 1M |
| Claude Opus 4.7 (Anthropic) | ~1753 | 64.3% (Pro) | ~91% | ~$15/$75 per 1M |
| Gemini 3.1 Pro (Google) | ~1314 | 54.2% (Pro) | 94.3% | $2/$12 per 1M |
Gemini 3.5 Flash represents a strategic repositioning of Google's 'Flash' tier from a budget model to a frontier-class agentic specialist. Launched at Google I/O 2026, it delivers state-of-the-art performance in multi-step tool use (MCP Atlas: 83.6%) and agentic coding (Terminal-Bench 2.1: 76.2%), often beating the previous Gemini 3.1 Pro while running approximately 4x faster. This makes it the default production backbone for coding agents, computer automation, and multimodal workflows within the Google ecosystem.
However, this upgrade comes with a significant 3x price increase over its predecessor, moving Flash into a mid-tier pricing bracket ($1.50/$9 per million tokens). While still 40% cheaper than Gemini 3.1 Pro, it is no longer the budget option. The model's key innovation is its thought preservation feature for multi-turn conversations and adjustable reasoning levels, but it has notable weaknesses: the 1M context window has poor recall at full length, it lacks computer-use support, and it trails Pro models on pure reasoning benchmarks like Humanity's Last Exam (40.2% vs. 44.4%).
In summary, Gemini 3.5 Flash is the premier choice for high-volume, speed-sensitive agentic applications on the Google stack, but developers must recalibrate expectations and budgets away from the old Flash tier's economics and evaluate whether Pro-tier models are still needed for the hardest reasoning tasks.
Sources
- Gemini 3.5 Flash — Google DeepMind (Official Page)
- Gemini 3.5 Flash benchmark scores — evals.report
- Gemini 3.5 Flash Benchmarks, Pricing & Speed — BenchLM.ai
- Google's Gemini 3.5 Flash Levels Up, Approaching Top Models But Raising Prices — The Batch (deeplearning.ai)
- Gemini 3.5 Flash, reviewed — benchr
- Gemini 3.5 Flash Review: Benchmarks, Price & API — buildfastwithai.com
- Gemini 3.1 Pro vs Gemini 3.5 Flash: Benchmarks, Pricing — BenchLM.ai
Analysis generated: 2026-07-17