Google Deep MindProprietary

Gemini 3.1 Flash-Lite

Compare this model

A high-performance multimodal model developed by Google DeepMind for integrated processing of text, images, and video.

Parameters

Undisclosed

Context Window

License

Proprietary

Release Date

2026-05-07

Japanese Language Capability

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API Pricing

API pricing for this model is not yet available

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        Release

        March 3, 2026 (preview)

        Cheapest, fastest Gemini 3-tier model; catalog GA 2026-05-07

        Context Window

        1,000,000 tokens

        64K-65.5K token output ceiling; text/image/audio/video/PDF input, text output

        Pricing

        $0.25 / $1.50 per 1M

        About 1/8 the cost of Gemini 3.1 Pro; no published cache discount

        Speed

        363 tokens/sec

        TTFT ~2.5x faster than Gemini 2.5 Flash

        GPQA Diamond

        86.9%

        vs GPT-5 mini 82.3%, Gemini 2.5 Flash 82.8% (no tools)

        Strengths

        • Best price-per-token in the Gemini 3 family at $0.25/$1.50, roughly 1/8 of Pro, for a 1M-token context and native multimodal input.
        • Throughput king: 363 tokens/sec and a time-to-first-token about 2.5x faster than Gemini 2.5 Flash, built for high-volume, latency-sensitive jobs.
        • Reasoning effort and reasoning budget controls let teams dial thinking down for classification/translation and up for harder tasks.
        • GPQA Diamond 86.9% and MMMU-Pro 76.8% hold their own against models that cost far more.

        Weaknesses

        • Deliberately not an agent model - Google scopes it to data-processing workloads and publishes no agentic benchmark scores.
        • Reasoning depth is capped versus 3.1 Pro; on the hardest long-context retrieval (MRCR 1M pointwise) it scores only ~12%, well behind Pro.
        • Output speed and capability still trail frontier reasoning models; treat it as a workhorse, not a flagship.
        • No open weights; only Google and approved partners serve it.

        Competitor Comparison

        ModelArenaGPQAPrice
        Gemini 3.1 Flash-Lite (this)143286.9%$0.25/$1.50
        GPT-5 mini-82.3%$1.00/$5.00
        Claude 4.5 Haiku-84.3%$1.00/$5.00
        Gemini 2.5 Flash-82.8%$0.30/$2.50

        Gemini 3.1 Flash-Lite is Google DeepMind's cheapest, fastest Gemini 3 model, distilled from Gemini 3 Pro and built for high-volume, latency-sensitive work - translation, classification, and document pipelines - at $0.25/$1.50 per million tokens with a 1M-token context.

        Analysis generated: 2026-09-08