A high-performance multimodal model developed by Google DeepMind for integrated processing of text, images, and video.
Parameters
Undisclosed
Context Window
License
Proprietary
Release Date
2026-05-07
Japanese Language Capability
Multilingual model with strong Japanese language processing capabilities.
API Pricing
API pricing for this model is not yet available
Strengths
Weaknesses
Use Cases
Deep Analysis
Release
March 3, 2026 (preview)
Cheapest, fastest Gemini 3-tier model; catalog GA 2026-05-07
Context Window
1,000,000 tokens
64K-65.5K token output ceiling; text/image/audio/video/PDF input, text output
Pricing
$0.25 / $1.50 per 1M
About 1/8 the cost of Gemini 3.1 Pro; no published cache discount
Speed
363 tokens/sec
TTFT ~2.5x faster than Gemini 2.5 Flash
GPQA Diamond
86.9%
vs GPT-5 mini 82.3%, Gemini 2.5 Flash 82.8% (no tools)
Strengths
- ・Best price-per-token in the Gemini 3 family at $0.25/$1.50, roughly 1/8 of Pro, for a 1M-token context and native multimodal input.
- ・Throughput king: 363 tokens/sec and a time-to-first-token about 2.5x faster than Gemini 2.5 Flash, built for high-volume, latency-sensitive jobs.
- ・Reasoning effort and reasoning budget controls let teams dial thinking down for classification/translation and up for harder tasks.
- ・GPQA Diamond 86.9% and MMMU-Pro 76.8% hold their own against models that cost far more.
Weaknesses
- ・Deliberately not an agent model - Google scopes it to data-processing workloads and publishes no agentic benchmark scores.
- ・Reasoning depth is capped versus 3.1 Pro; on the hardest long-context retrieval (MRCR 1M pointwise) it scores only ~12%, well behind Pro.
- ・Output speed and capability still trail frontier reasoning models; treat it as a workhorse, not a flagship.
- ・No open weights; only Google and approved partners serve it.
Competitor Comparison
| Model | Arena | GPQA | Price |
|---|---|---|---|
| Gemini 3.1 Flash-Lite (this) | 1432 | 86.9% | $0.25/$1.50 |
| GPT-5 mini | - | 82.3% | $1.00/$5.00 |
| Claude 4.5 Haiku | - | 84.3% | $1.00/$5.00 |
| Gemini 2.5 Flash | - | 82.8% | $0.30/$2.50 |
Gemini 3.1 Flash-Lite is Google DeepMind's cheapest, fastest Gemini 3 model, distilled from Gemini 3 Pro and built for high-volume, latency-sensitive work - translation, classification, and document pipelines - at $0.25/$1.50 per million tokens with a 1M-token context.
Sources
Analysis generated: 2026-09-08