AlibabaProprietary

Qwen3.7-Max-Preview

Compare this model

The latest agent-specialized model developed by Alibaba. It enables advanced agent workflows and reasoning capabilities.

Parameters

Undisclosed

Context Window

License

Proprietary

Release Date

2026-05-20

Japanese Language Capability

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API Pricing

Input Price (per 1M tokens)

$2.5

Output Price (per 1M tokens)

$

Billing Mode: standard

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        Artificial Analysis Intelligence Index

        56.6

        #5 overall out of 218+ ranked models; highest Chinese model; +4.8 over Qwen3.6 Max Preview

        Arena AI Elo

        ~1,489 (#14 overall)

        #7 Math, #9 Expert Prompts, #9 Software & IT; WebDev Arena #4 with 1,541 Elo

        SWE-Bench Verified

        80.4%

        SWE-Bench Pro: 60.6%; SWE Multilingual: 78.3%; leading domestic models

        GPQA Diamond

        92.4%

        PhD-level reasoning; above Claude Opus 4.6 (91.3%), below GPT-5.5 (93.6%)

        Input / Output Price

        $2.50 / $7.50 per 1M tokens

        Cached input: $0.25/1M (90% discount); ~half of Claude Opus 4.7 pricing

        Context Window

        1M tokens

        4× increase from 256K on Qwen3.6 Max Preview; max output 64K–66K tokens

        Strengths

        • Best price-to-intelligence ratio among frontier models — matches or beats Claude Opus 4.6 on multiple benchmarks at roughly one-third the output cost
        • Exceptional long-horizon agentic capability with 1M-token context and demonstrated 35-hour autonomous execution with 1,000+ tool calls
        • Industry-leading abstinence behavior: lowest hallucination rate (22.9%) among frontier models on AA-Omniscience, with strong GPQA Diamond (92.4%) and HMMT (97.1%) scores

        Weaknesses

        • Text-only with zero vision capability — completely disqualifies any workflow requiring image understanding, UI screenshots, or computer-use agents
        • High verbosity in reasoning mode generates ~97M output tokens vs. 26M median on benchmarks, significantly increasing latency (10–45s thinking overhead) and per-task API costs
        • No open weights available; proprietary API-only access limits local deployment, fine-tuning, and offline use cases compared to earlier Qwen releases

        Competitor Comparison

        ModelArenaSWEGPQAPrice
        GPT-5.5 (OpenAI)~1,550+N/A (published)93.6%$5.00/$30.00
        Claude Opus 4.7 (Anthropic)Top 564.3% (Pro)~91–92%$5.00/$25.00
        Gemini 3.1 Pro Preview (Google)Top 10N/AN/AN/A (preview)
        Gemini 3.5 Flash (Google)Top 10N/AN/A$1.50/$9.00
        DeepSeek V4 ProBelow Qwen3.7-MaxN/AN/A$1.74/$3.48

        Qwen3.7-Max-Preview is Alibaba's flagship proprietary reasoning model, released on May 20–21, 2026, at the Alibaba Cloud Summit in Hangzhou. Positioned as an 'agent foundation' rather than a general-purpose chat assistant, it targets long-horizon autonomous coding, scientific reasoning, and office workflow automation. The model scored 56.6 on the Artificial Analysis Intelligence Index (#5 overall), making it the highest-ranked Chinese model on that leaderboard. It holds a 1M-token context window (4× the previous generation), delivers 92.4% on GPQA Diamond, 80.4% on SWE-Bench Verified, and demonstrated a 35-hour autonomous coding run on Alibaba's custom Zhenwu M890 accelerator in internal testing.

        The model occupies a distinctive sweet spot in the market: it approaches the raw intelligence of Claude Opus 4.7 and Gemini 3.1 Pro Preview while costing roughly half on input tokens and a third on output tokens. Its aggressive prompt caching ($0.25/1M cached input — a 90% discount) makes it particularly attractive for agent workflows that re-read long contexts across hundreds of turns. The model natively supports both OpenAI and Anthropic API specifications, enabling drop-in integration with tools like Claude Code. However, it is text-only (no vision), generates very verbose reasoning traces that inflate token costs, and has a notably high abstention rate (48% attempt rate on AA-Omniscience), which reduces hallucinations but limits usefulness on knowledge-recall tasks.

        Alibaba's release marks a strategic shift toward closed-weight, revenue-generating models at the top tier while maintaining open weights for lower-tier variants (announced 27B dense and 35B MoE models have no confirmed release date). The one-month cadence from Qwen3.6 to Qwen3.7 signals an aggressive development pace. For engineering teams, Qwen3.7-Max represents the strongest cost-optimized frontier option for text-only agentic and coding workloads, though teams requiring vision, absolute peak intelligence, or low-latency chat should look elsewhere.

        Analysis generated: 2026-07-17