The latest agent-specialized model developed by Alibaba. It enables advanced agent workflows and reasoning capabilities.
Parameters
Undisclosed
Context Window
License
Proprietary
Release Date
2026-05-20
Japanese Language Capability
Multilingual model with strong Japanese language processing capabilities.
API Pricing
Input Price (per 1M tokens)
$2.5
Output Price (per 1M tokens)
$
Billing Mode: standard
Strengths
Weaknesses
Use Cases
Deep Analysis
Artificial Analysis Intelligence Index
56.6
#5 overall out of 218+ ranked models; highest Chinese model; +4.8 over Qwen3.6 Max Preview
Arena AI Elo
~1,489 (#14 overall)
#7 Math, #9 Expert Prompts, #9 Software & IT; WebDev Arena #4 with 1,541 Elo
SWE-Bench Verified
80.4%
SWE-Bench Pro: 60.6%; SWE Multilingual: 78.3%; leading domestic models
GPQA Diamond
92.4%
PhD-level reasoning; above Claude Opus 4.6 (91.3%), below GPT-5.5 (93.6%)
Input / Output Price
$2.50 / $7.50 per 1M tokens
Cached input: $0.25/1M (90% discount); ~half of Claude Opus 4.7 pricing
Context Window
1M tokens
4× increase from 256K on Qwen3.6 Max Preview; max output 64K–66K tokens
Strengths
- ・Best price-to-intelligence ratio among frontier models — matches or beats Claude Opus 4.6 on multiple benchmarks at roughly one-third the output cost
- ・Exceptional long-horizon agentic capability with 1M-token context and demonstrated 35-hour autonomous execution with 1,000+ tool calls
- ・Industry-leading abstinence behavior: lowest hallucination rate (22.9%) among frontier models on AA-Omniscience, with strong GPQA Diamond (92.4%) and HMMT (97.1%) scores
Weaknesses
- ・Text-only with zero vision capability — completely disqualifies any workflow requiring image understanding, UI screenshots, or computer-use agents
- ・High verbosity in reasoning mode generates ~97M output tokens vs. 26M median on benchmarks, significantly increasing latency (10–45s thinking overhead) and per-task API costs
- ・No open weights available; proprietary API-only access limits local deployment, fine-tuning, and offline use cases compared to earlier Qwen releases
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| GPT-5.5 (OpenAI) | ~1,550+ | N/A (published) | 93.6% | $5.00/$30.00 |
| Claude Opus 4.7 (Anthropic) | Top 5 | 64.3% (Pro) | ~91–92% | $5.00/$25.00 |
| Gemini 3.1 Pro Preview (Google) | Top 10 | N/A | N/A | N/A (preview) |
| Gemini 3.5 Flash (Google) | Top 10 | N/A | N/A | $1.50/$9.00 |
| DeepSeek V4 Pro | Below Qwen3.7-Max | N/A | N/A | $1.74/$3.48 |
Qwen3.7-Max-Preview is Alibaba's flagship proprietary reasoning model, released on May 20–21, 2026, at the Alibaba Cloud Summit in Hangzhou. Positioned as an 'agent foundation' rather than a general-purpose chat assistant, it targets long-horizon autonomous coding, scientific reasoning, and office workflow automation. The model scored 56.6 on the Artificial Analysis Intelligence Index (#5 overall), making it the highest-ranked Chinese model on that leaderboard. It holds a 1M-token context window (4× the previous generation), delivers 92.4% on GPQA Diamond, 80.4% on SWE-Bench Verified, and demonstrated a 35-hour autonomous coding run on Alibaba's custom Zhenwu M890 accelerator in internal testing.
The model occupies a distinctive sweet spot in the market: it approaches the raw intelligence of Claude Opus 4.7 and Gemini 3.1 Pro Preview while costing roughly half on input tokens and a third on output tokens. Its aggressive prompt caching ($0.25/1M cached input — a 90% discount) makes it particularly attractive for agent workflows that re-read long contexts across hundreds of turns. The model natively supports both OpenAI and Anthropic API specifications, enabling drop-in integration with tools like Claude Code. However, it is text-only (no vision), generates very verbose reasoning traces that inflate token costs, and has a notably high abstention rate (48% attempt rate on AA-Omniscience), which reduces hallucinations but limits usefulness on knowledge-recall tasks.
Alibaba's release marks a strategic shift toward closed-weight, revenue-generating models at the top tier while maintaining open weights for lower-tier variants (announced 27B dense and 35B MoE models have no confirmed release date). The one-month cadence from Qwen3.6 to Qwen3.7 signals an aggressive development pace. For engineering teams, Qwen3.7-Max represents the strongest cost-optimized frontier option for text-only agentic and coding workloads, though teams requiring vision, absolute peak intelligence, or low-latency chat should look elsewhere.
Sources
- Alibaba's Latest Proprietary Model, Qwen3.7-Max, Challenges U.S. Rivals — The Batch (DeepLearning.AI)
- Qwen 3.7 Max Preview: Benchmarks, Rankings & Open Weights — Bytepith
- Qwen Introduces Qwen3.7-Max: A Reasoning Agent Model With a 1M-Token Context Window — MarkTechPost
- Qwen3.7-Max Performance Benchmarks: Terminal Bench Leader, SWE Scores & 35-Hour T-Head PPU Run — TheRouter.ai
- Qwen3.7-Max Review 2026: Benchmarks, Pricing, Verdict — FelloAI
- How to Use Qwen3.7-Max: Qwen Chat, Pricing & Benchmarks — 7minAI
- Qwen 3.6 Max (preview) vs Qwen3.7 Max: Benchmarks, Pricing, Speed — BenchLM.ai
- Qwen 3.7 Max Review: The $2.50 Model That's Beating Claude Opus 4.6 at Half the Price — MEFAI
- Qwen3.6 Max Preview vs Qwen3.7 Max: Pricing, Benchmarks & API Compared — OminiGate
- 阿里云千问大模型 Qwen3.7-Max-Preview 首发亮相 Arena AI — IT之家
- Qwen3.7 Official Blog Post — Qwen (Alibaba)
Analysis generated: 2026-07-17