Google Deep Mindのfoundationモデル。
パラメータ
非公開
コンテキスト長
1M
ライセンス
プロプライエタリ
リリース日
2026-07-21
API料金
入力料金(1Mトークンあたり)
$0.3
出力料金(1Mトークンあたり)
$2.5
課金モード: standard
強み
弱み
活用例
深度分析
Context Window
1,048,576 tokens
1M multimodal (text, image, audio, video, PDF)
Artificial Analysis Intelligence Index
36
#12 of 152 models; top of its class
Input Price
$0.30 / 1M tokens
Output: $2.50 / 1M; batch 50% off
Throughput
~490 tok/s
2nd-fastest measured on Artificial Analysis
SWE-Bench Pro
54.2%
vs Gemini 3.1 Flash-Lite 38%
Terminal-Bench 2.1
54.0%
vs Gemini 3.1 Flash-Lite 31%
OSWorld-Verified
74.0%
computer use; ~2x predecessor
強み
- ・Cheapest paid Gemini 3 tier at $0.30/$2.50 per 1M tokens, undercutting GPT-5.4 mini and Claude Haiku 4.5 on input price
- ・1M-token multimodal context with thinking, function calling, structured outputs, code execution and search grounding
- ・Near-top speed (~490 tok/s) makes it ideal for high-volume, latency-sensitive production traffic
弱み
- ・Loses to GPT-5.4 mini on sustained multi-step reasoning (Terminal-Bench 2.1, GDPval-AA v2)
- ・No computer use, image/audio generation, or Live API — a text-in/text-out engine for volume
- ・Time-to-first-token around 6s (includes thinking), so better suited to background jobs than snappy live chat
競合比較
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| Gemini 3.5 Flash-Lite | N/A | 54.2% | N/A | $0.30 / $2.50 per 1M |
| GPT-5.4 mini | N/A | N/A | N/A | $0.75 / $4.50 per 1M |
| Claude Haiku 4.5 | N/A | N/A | N/A | $1.00 / $5.00 per 1M |
| Gemini 3.1 Flash-Lite | N/A | 38% | N/A | prior budget tier |
Gemini 3.5 Flash-Lite is Google DeepMind's cheapest Gemini 3 model, released July 21, 2026 alongside Gemini 3.6 Flash for high-volume, latency-sensitive workloads such as translation, classification, and tagging. It carries a 1,048,576-token multimodal context (text, image, audio, video, PDF) at just $0.30 per million input and $2.50 per million output tokens, with a flat 50% discount in batch mode for non-real-time jobs.
Google positions Flash-Lite as the layer beneath Gemini 3.6 Flash: 3.6 Flash for complex agent work, 3.5 Flash-Lite for the high-volume, latency-sensitive jobs that make up most production traffic — agentic search, document processing, translation, classification. On Artificial Analysis it lands an Intelligence Index of 36 (#12 of 152 models, well above the class average) while running at roughly 490 output tokens per second, the second-fastest model the service has measured. The generation jump over its March 2026 preview predecessor, Gemini 3.1 Flash-Lite, is the headline: SWE-Bench Pro nearly closes a 16-point gap (54.2% vs 38%), OSWorld-Verified climbs almost 20 points to 74.0%, and Terminal-Bench 2.1 doubles to 54.0%.
出典
分析生成日: 2026-09-02