The latest agent-specialized model from Alibaba, enabling advanced agent workflows and reasoning capabilities.
Parameters
Undisclosed
Context Window
256K
License
Proprietary
Release Date
2026-05-20
Japanese Language Capability
Multilingual model with strong Japanese language processing capabilities.
API Pricing
API pricing for this model is not yet available
Strengths
Weaknesses
Use Cases
Deep Analysis
Arena Elo
1460
#15 overall, #12 in coding (BenchLM provisional leaderboard)
SWE-Bench Verified
77.7%
vs Qwen3.7-Max: 80.4%
ScreenSpot Pro (GUI grounding)
79.0
Frontier-tier for computer use agents
Input Price
$0.40/1M tokens
~6x cheaper than Qwen3.7-Max
Context Window
1M tokens
Shared across text, image, and video
Terminal-Bench 2.0
70.3%
Slightly ahead of Qwen3.7-Max (69.7%)
Strengths
- ・Exceptional GUI grounding (ScreenSpot Pro 79.0) for computer use and browser automation
- ・Multimodal interactive hybrid agent unifying vision, code generation, and GUI/CLI operation
- ・Budget-tier pricing with $0.40/1M input tokens while maintaining near-Max text performance
Weaknesses
- ・Proprietary API-only model with no open weights for self-hosting or customization
- ・Slightly trails text-only Qwen3.7-Max on pure-text and math benchmarks (e.g., SWE-Bench Verified 77.7% vs 80.4%)
- ・Visual tokens consume shared 1M context budget, making heavy image/video workloads expensive
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| Qwen3.7-Max | 1475 | 80.4% | 92.4% | $2.50/$7.50 |
| Claude Sonnet 4.6 | 1480 | ~77% (estimated) | ~90% (estimated) | $3/$15 |
| DeepSeek V4 Pro (Max) | 1485 | ~80% (estimated) | ~91% (estimated) | $1.50/$4.50 |
Qwen3.7-Plus is Alibaba's flagship multimodal agent model, released in June 2026 as the vision-capable counterpart to the text-only Qwen3.7-Max. It builds on the Qwen3.7 backbone to deliver frontier-tier GUI grounding (ScreenSpot Pro 79.0) and hybrid GUI+CLI agent capabilities, enabling end-to-end automation of tasks like browser navigation, UI recreation, and software development from screenshots. The model uniquely combines vision perception, code generation, and tool use within a single agent loop, allowing it to "see, think, write, act, and verify" in closed-loop execution.
Positioned as a "budget multimodal" tier, Plus is roughly six times cheaper than Qwen3.7-Max on input tokens while matching or slightly trailing it on pure-text benchmarks. Its 1M-token context window and 35-hour autonomous run ceiling support long-horizon agent workflows, and it generalizes across agent frameworks like Claude Code, OpenClaw, and Qwen Code. However, it remains API-only with no open weights, which may limit adoption in regulated or air-gapped environments. The model represents Alibaba's strategic move to compete with Western frontier models in the high-stakes GUI agent and computer-use space, offering a cost-effective alternative for teams building screen-aware automation.
Sources
- Qwen3.7-Plus: Multimodal Agent Intelligence - Alibaba Cloud Community
- Qwen3.7 Plus Benchmarks, Pricing & Speed (July 2026) | BenchLM.ai
- Qwen 3.7 Plus vs Max: which Qwen 3.7 model should you use?
- Qwen 3.7 Plus: Alibaba's multimodal agent model, benchmarks and pricing
- Qwen3.7 Max vs Qwen3.7 Plus: Benchmarks, Pricing, Speed (July 2026) | BenchLM.ai
- Qwen3.7-Plus Review: Alibaba's GUI Agent, Tested
- Qwen3.7-Plus:想得深,看得懂,做得到-阿里云开发者社区
- Qwen3.7-Plus上线千问云,多模态智能体能力再升级!-阿里云开发者社区
Analysis generated: 2026-07-17