Gemini 3.5 Flash-Lite is Google DeepMind's cheapest Gemini 3 model, released July 21, 2026 for high-volume, latency-sensitive workloads like translation, classification, and tagging. It serves a 1,000,000-token multimodal context (text, image, audio, video, PDF) at just $0.30 per million input and $2.50 per million output tokens. It beats Gemini 3.1 Flash-Lite on agentic coding (SWE-Bench Pro 54.2%), computer use (OSWorld-Verified 74.0%), and terminal tasks (Terminal-Bench 2.1 54%), and streams at roughly 350 tokens per second on the Artificial Analysis Index.
Parameters
Undisclosed
Context Window
1M
License
Proprietary
Release Date
2026-07-21
API Pricing
Input Price (per 1M tokens)
$0.3
Output Price (per 1M tokens)
$2.5
Billing Mode: standard
Strengths
- •Optimized for speed and efficiency in inference.
- •Good balance of performance and reduced computational overhead.
- •Effective for a broad range of general language tasks.
Weaknesses
- •May not match the peak accuracy of larger, more complex models.
- •Capabilities in highly specialized or nuanced domains could be limited.
- •Performance scales with efficient resource allocation.
Use Cases
- •High-volume, latency-sensitive applications like chatbots.
- •Quick prototyping and development of AI-powered features.
- •Efficient processing of large text datasets for summarization or extraction.
Deep Analysis
Context Window
1,048,576 tokens
1M multimodal (text, image, audio, video, PDF)
Artificial Analysis Intelligence Index
36
#12 of 152 models; top of its class
Input Price
$0.30 / 1M tokens
Output: $2.50 / 1M; batch 50% off
Throughput
~490 tok/s
2nd-fastest measured on Artificial Analysis
SWE-Bench Pro
54.2%
vs Gemini 3.1 Flash-Lite 38%
Terminal-Bench 2.1
54.0%
vs Gemini 3.1 Flash-Lite 31%
OSWorld-Verified
74.0%
computer use; ~2x predecessor
Strengths
- ・Cheapest paid Gemini 3 tier at $0.30/$2.50 per 1M tokens, undercutting GPT-5.4 mini and Claude Haiku 4.5 on input price
- ・1M-token multimodal context with thinking, function calling, structured outputs, code execution and search grounding
- ・Near-top speed (~490 tok/s) makes it ideal for high-volume, latency-sensitive production traffic
Weaknesses
- ・Loses to GPT-5.4 mini on sustained multi-step reasoning (Terminal-Bench 2.1, GDPval-AA v2)
- ・No computer use, image/audio generation, or Live API — a text-in/text-out engine for volume
- ・Time-to-first-token around 6s (includes thinking), so better suited to background jobs than snappy live chat
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | N/A | 54.2% | N/A | $0.30 / $2.50 per 1M |
| GPT-5.4 mini | N/A | N/A | N/A | $0.75 / $4.50 per 1M |
| Claude Haiku 4.5 | N/A | N/A | N/A | $1.00 / $5.00 per 1M |
| Gemini 3.1 Flash-Lite | N/A | 38% | N/A | prior budget tier |
Gemini 3.5 Flash-Lite is Google DeepMind's cheapest Gemini 3 model, released July 21, 2026 alongside Gemini 3.6 Flash for high-volume, latency-sensitive workloads such as translation, classification, and tagging. It carries a 1,048,576-token multimodal context (text, image, audio, video, PDF) at just $0.30 per million input and $2.50 per million output tokens, with a flat 50% discount in batch mode for non-real-time jobs.
Google positions Flash-Lite as the layer beneath Gemini 3.6 Flash: 3.6 Flash for complex agent work, 3.5 Flash-Lite for the high-volume, latency-sensitive jobs that make up most production traffic — agentic search, document processing, translation, classification. On Artificial Analysis it lands an Intelligence Index of 36 (#12 of 152 models, well above the class average) while running at roughly 490 output tokens per second, the second-fastest model the service has measured. The generation jump over its March 2026 preview predecessor, Gemini 3.1 Flash-Lite, is the headline: SWE-Bench Pro nearly closes a 16-point gap (54.2% vs 38%), OSWorld-Verified climbs almost 20 points to 74.0%, and Terminal-Bench 2.1 doubles to 54.0%.
Sources
Analysis generated: 2026-09-02