Google ResearchProprietary

Gemini-SQL2

Compare this model

An SQL-specialized model developed by Google Research, optimized for database query generation.

Parameters

Undisclosed

Context Window

License

Proprietary

Release Date

2026-06-03

Japanese Language Capability

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API Pricing

API pricing for this model is not yet available

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        BIRD Execution Accuracy

        80.04%

        #1 on single-model leaderboard

        Gap to Human Performance

        12.92 pts

        Human: 92.96%

        Lead over GPT-5.5-xhigh

        +7.2 pts

        GPT-5.5-xhigh: ~72.8%

        Lead over Claude Opus 4.6

        +9.1 pts

        Claude Opus 4.6: ~70.9%

        Base Model

        Gemini 3.1 Pro

        Specialized post-training on top of general model

        Public API Status

        Not released

        No model card, API, or technical paper as of Jun 2026

        Strengths

        • Clear state-of-the-art lead on BIRD, the hardest execution-verified text-to-SQL benchmark, beating all competitors by 3–10 points
        • Built on Gemini 3.1 Pro's strong general reasoning, meaning it handles complex joins, window functions, and ambiguous schemas better than specialized 32B models
        • Planned native integration into BigQuery Studio, AlloyDB AI, and Cloud SQL Studio gives Google a structural distribution advantage no competitor has

        Weaknesses

        • No public API, model card, or technical report released — the 80.04% claim cannot be independently audited or reproduced
        • ~20% error rate (1 in 5 queries) still makes it unsuitable for unsupervised production use; human review remains mandatory
        • Benchmark performance on curated BIRD schemas may not translate to messy enterprise databases with undocumented joins and legacy naming conventions

        Competitor Comparison

        ModelPrice
        Gemini-SQL2 (Google)Not public
        Gemini-SQL (Google)Not public
        Q-SQL (AWS)Not public
        GPT-5.5-xhigh (OpenAI)API-based
        Claude Opus 4.6 (Anthropic)API-based

        Gemini-SQL2 is Google Research's specialized text-to-SQL capability built on Gemini 3.1 Pro, announced on June 12, 2026. It achieved 80.04% execution accuracy on the BIRD benchmark's single-model track—the first system to clear 80%—establishing a decisive lead over OpenAI's GPT-5.5-xhigh (~72.8%), Anthropic's Claude Opus 4.6 (~70.9%), and specialized models from AWS, Databricks, Tencent, Alibaba, and Snowflake. The system translates natural language questions into execution-ready SQL, meaning queries must not only parse syntactically but also run against real databases and return correct results.

        The positioning is strategic rather than incidental. Google now holds the top two named positions on the BIRD leaderboard (Gemini-SQL2 and its predecessor Gemini-SQL at ~77.2%), and the company has signaled planned integration into BigQuery Studio, AlloyDB AI, and Cloud SQL Studio. This gives Google a structural advantage: it owns both the model and the databases where it will ship. No competitor has a comparable first-party distribution path for enterprise text-to-SQL.

        However, significant caveats remain. No public API, model card, or technical paper has been released, meaning the 80.04% claim is unverified by the research community. The 12.92-point gap to human performance (92.96%) translates to roughly a 1-in-5 query failure rate—meaning human review is still mandatory for production use. Multiple practitioners have noted that BIRD's curated schemas may not represent the messiness of real enterprise databases, and Google has provided no evidence of production-scale deployment testing.

        Analysis generated: 2026-07-17