Google Deep MindProprietary

Gemini 3.5 Flash-Lite

Compare this model

Gemini 3.5 Flash-Lite is Google DeepMind's cheapest Gemini 3 model, released July 21, 2026 for high-volume, latency-sensitive workloads like translation, classification, and tagging. It serves a 1,000,000-token multimodal context (text, image, audio, video, PDF) at just $0.30 per million input and $2.50 per million output tokens. It beats Gemini 3.1 Flash-Lite on agentic coding (SWE-Bench Pro 54.2%), computer use (OSWorld-Verified 74.0%), and terminal tasks (Terminal-Bench 2.1 54%), and streams at roughly 350 tokens per second on the Artificial Analysis Index.

Parameters

Undisclosed

Context Window

1M

License

Proprietary

Release Date

2026-07-21

API Pricing

Input Price (per 1M tokens)

$0.3

Output Price (per 1M tokens)

$2.5

Billing Mode: standard

Strengths

  • Optimized for speed and efficiency in inference.
  • Good balance of performance and reduced computational overhead.
  • Effective for a broad range of general language tasks.

Weaknesses

  • May not match the peak accuracy of larger, more complex models.
  • Capabilities in highly specialized or nuanced domains could be limited.
  • Performance scales with efficient resource allocation.

Use Cases

  • High-volume, latency-sensitive applications like chatbots.
  • Quick prototyping and development of AI-powered features.
  • Efficient processing of large text datasets for summarization or extraction.

Deep Analysis

Context Window

1,048,576 tokens

1M multimodal (text, image, audio, video, PDF)

Artificial Analysis Intelligence Index

36

#12 of 152 models; top of its class

Input Price

$0.30 / 1M tokens

Output: $2.50 / 1M; batch 50% off

Throughput

~490 tok/s

2nd-fastest measured on Artificial Analysis

SWE-Bench Pro

54.2%

vs Gemini 3.1 Flash-Lite 38%

Terminal-Bench 2.1

54.0%

vs Gemini 3.1 Flash-Lite 31%

OSWorld-Verified

74.0%

computer use; ~2x predecessor

Strengths

  • Cheapest paid Gemini 3 tier at $0.30/$2.50 per 1M tokens, undercutting GPT-5.4 mini and Claude Haiku 4.5 on input price
  • 1M-token multimodal context with thinking, function calling, structured outputs, code execution and search grounding
  • Near-top speed (~490 tok/s) makes it ideal for high-volume, latency-sensitive production traffic

Weaknesses

  • Loses to GPT-5.4 mini on sustained multi-step reasoning (Terminal-Bench 2.1, GDPval-AA v2)
  • No computer use, image/audio generation, or Live API — a text-in/text-out engine for volume
  • Time-to-first-token around 6s (includes thinking), so better suited to background jobs than snappy live chat

Competitor Comparison

ModelArenaSWEGPQAPrice
Gemini 3.5 Flash-LiteN/A54.2%N/A$0.30 / $2.50 per 1M
GPT-5.4 miniN/AN/AN/A$0.75 / $4.50 per 1M
Claude Haiku 4.5N/AN/AN/A$1.00 / $5.00 per 1M
Gemini 3.1 Flash-LiteN/A38%N/Aprior budget tier

Gemini 3.5 Flash-Lite is Google DeepMind's cheapest Gemini 3 model, released July 21, 2026 alongside Gemini 3.6 Flash for high-volume, latency-sensitive workloads such as translation, classification, and tagging. It carries a 1,048,576-token multimodal context (text, image, audio, video, PDF) at just $0.30 per million input and $2.50 per million output tokens, with a flat 50% discount in batch mode for non-real-time jobs.

Google positions Flash-Lite as the layer beneath Gemini 3.6 Flash: 3.6 Flash for complex agent work, 3.5 Flash-Lite for the high-volume, latency-sensitive jobs that make up most production traffic — agentic search, document processing, translation, classification. On Artificial Analysis it lands an Intelligence Index of 36 (#12 of 152 models, well above the class average) while running at roughly 490 output tokens per second, the second-fastest model the service has measured. The generation jump over its March 2026 preview predecessor, Gemini 3.1 Flash-Lite, is the headline: SWE-Bench Pro nearly closes a 16-point gap (54.2% vs 38%), OSWorld-Verified climbs almost 20 points to 74.0%, and Terminal-Bench 2.1 doubles to 54.0%.

Analysis generated: 2026-09-02