Qwen3.8-Max-0902, deployed September 2, 2026 via QwenCloud, is a post-training refresh of Alibaba's Qwen3.8-Max. The underlying 2.4T-parameter, 95B-active MoE architecture is unchanged from the August 12 checkpoint, but the model climbed to #1 on Code Arena's blind WebDev leaderboard at 1,691 points and lifted TerminalBench 3.0 from 11.3 to 29.0. Context (1M tokens) and $2/$6 pricing are unchanged; availability is API-only across OpenAI/Anthropic-compatible endpoints in Beijing, Singapore and Virginia. No new open-weights drop — the downloadable Qwen3.8-Max remains the August 12 release.
Parameters
2.4T total / 95B active (MoE)
Context Window
1M (991K input / 131K output / 262K reasoning budget)
License
Closed (API-only via QwenCloud)
Release Date
2026-09-02
API Pricing
Input Price (per 1M tokens)
$2
Output Price (per 1M tokens)
$6
Billing Mode: blended $5/MTok, cache read $0.17
Strengths
- •Code Arena WebDev #1 at 1,691 — first Chinese model to top a blind, real-user-voted benchmark
- •TerminalBench 3.0 more than doubled (11.3 → 29.0); QwenSWEbench V2 hit 70.0
- •OpenAI- and Anthropic-compatible APIs; direct integration with Claude Code, Codex, Qoder and Qwen Code
- •Blended $5/MTok — best price-to-score ratio on the Code Arena Pareto frontier
Weaknesses
- •No open weights; API-only and routed through Chinese infrastructure under Chinese jurisdiction
- •The 3-point WebDev lead over Opus 5 Max sits inside the confidence interval (±19 points)
- •Still trails Opus 5 Max on TerminalBench, DeepSWE and agent-coordination benchmarks
- •The '0902' suffix is a date stamp, not a version — behavior under the same model string can shift silently
Use Cases
- •Front-end and repository-level coding where the Code Arena WebDev lead matters
- •Long-horizon office automation (CoWorkBench-style workflows)
- •Cost-sensitive code review and refactoring at $6 per million output tokens
- •Drop-in OpenAI-compatible backend for Claude Code, Codex and Qoder
Deep Analysis
Context Window
1M tokens
991K max input; 131K max output; 262K reasoning budget; text/image/video input
API Price
$2 / $6 per 1M
Input/output per million tokens; blended $5/MTok; cache read $0.17
Code Arena WebDev
#1 at 1,691
First Chinese model to top the blind leaderboard; 3 above Claude Opus 5 Max
TerminalBench 3.0
29.0
Up from 11.3 on the prior Qwen3.8-Max checkpoint (2.6x)
Parameters
2.4T total / 95B active
MoE; same architecture as Qwen3.8-Max, post-training refresh
License
Closed (API-only)
No new open-weights drop; checkpoint remains the Aug 12 release
Strengths
- ・Genuine, targeted coding lift: TerminalBench 3.0 more than doubled (11.3 to 29.0) and QwenSWEbench V2 hit 70.0, all at unchanged $2/$6 pricing.
- ・Top of Code Arena WebDev at 1,691 — the first Chinese model to lead that blind, real-user-voted benchmark.
- ・Best price-to-score on the leaderboard's Pareto frontier at a blended $5/MTok — roughly 1/5 the input cost of Fable 5.1.
- ・Drop-in integrations: OpenAI- and Anthropic-compatible APIs plus Claude Code, Codex, Qoder and Qwen Code clients.
Weaknesses
- ・Not an open-weights release — API-only on Chinese infrastructure, so every call answers to Chinese jurisdiction.
- ・The 3-point WebDev lead over Opus 5 Max is inside the confidence interval (plus/minus 19 points, ~1,389 votes vs Opus's 10,000+); don't over-read it.
- ・Still trails Opus 5 on most of Qwen's own comparison rows: TerminalBench, DeepSWE, CoWorkBench, Toolathlon.
- ・The 0902 is a date-stamped snapshot, not a version — behavior can shift under the same model string without notice; pin your snapshot.
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| Qwen3.8-Max-0902 | #1 WebDev 1,691 | 70.0 (QwenSWE V2) | N/A | $2/$6 |
| Claude Opus 5 Max | WebDev 1,688 | 70.0 (QwenSWE V2) | N/A | $10/$50 |
| Kimi K3 Max | WebDev 1,674 | N/A | N/A | N/A |
| GLM-5.3 Flash | 57 (AA Index) | 63.4 (Z.ai) | N/A | $0.15/$0.50 |
Qwen3.8-Max-0902, shipped September 2, 2026, is a post-training refresh of Alibaba's 2.4T-parameter flagship — same 1M context, same $2/$6 pricing — that climbed to #1 on Code Arena's WebDev leaderboard at 1,691 points. It is the first Chinese-developed model to top that blind, real-user-voted benchmark, and it did so with a 22-point gain over its own prior checkpoint. The story is not new architecture; it is that a large share of perceived model quality is now post-training, delivered at unchanged price.
Sources
Analysis generated: 2026-09-06