GLM-5.3 Flash is Zhipu AI (Z.ai)'s 320B-parameter MoE (18B active) open-weight model, released August 26, 2026. It debuted anonymously as 'Ox-Alpha' and held OpenRouter's weekly top slot before Zhipu confirmed authorship on August 26. A sparse + linear attention hybrid lets it serve a 1.31M-token context cheaply, and it is the first native-multimodal GLM-5 release. Priced at $0.15/$0.50 per million tokens — about 1/40 of Claude Opus 4.8 — it ties Opus 4.8 on the AA Intelligence Index (57) and tops it on the vendor's own Code Bench (63.4). MIT-licensed.
Parameters
320B total / 18B active (MoE, sparse + linear attention hybrid)
Context Window
1.31M (131K max output)
License
MIT
Release Date
2026-08-26
API Pricing
Input Price (per 1M tokens)
$0.15
Output Price (per 1M tokens)
$0.5
Billing Mode: launch promo input $0.075 through 2026-09-09, then standard $0.15/$0.50
Strengths
- •Priced ~1/40 of Claude Opus 4.8 with comparable Artificial Analysis intelligence
- •First native-multimodal GLM-5 (image + video input)
- •Sparse + linear attention hybrid cuts 1M-context serving cost
- •MIT license: fine-tune, resell, or self-host with no constraints
Weaknesses
- •4-bit self-host still needs ~190GB of combined RAM/VRAM — not a single-GPU model
- •The 100k-domestic-chip serving claim is unverified and reads as marketing
- •AA Index 57 is frontier-adjacent, not frontier-leading on the hardest reasoning and agentic tests
- •Launch promo ($0.075 input) expires Sep 9; the real baseline is $0.15/$0.50
Use Cases
- •High-volume coding, classification, extraction and summarization where cost dominates
- •Document + image multimodal workflows without a second vision model
- •Replacing or downgrading Claude Opus subscriptions at ~1/40 the cost
- •Self-hosted fine-tuning research where the MIT license matters
Deep Analysis
Context Window
1.31M tokens
131K max output; native image + video input; first native-multimodal GLM-5 release
API Price
$0.15 / $0.50 per 1M
Input/output per million tokens; launch promo $0.075 input through Sep 9
AA Intelligence Index
57
Ties Claude Opus 4.8 on Artificial Analysis composite index
Z.ai Code Bench
63.4
Ahead of Claude Opus 4.8 (58.0) and DeepSeek-V4-Vision-Exp (59.3)
License
MIT (open weights)
Weights published on Hugging Face Aug 26, 2026
Parameters
320B total / 18B active
MoE; sparse + linear attention hybrid; first open frontier model with that mix
Strengths
- ・Price-performance is the headline: ~1/40 the cost of Claude Opus 4.8 at comparable coding quality on Z.ai Code Bench.
- ・First open-weight GLM-5 model with native multimodal input — can judge when to look and steer with visual feedback.
- ・Sparse + linear attention hybrid cuts 1M-context serving cost sharply; runs on far less hardware than the parameter count suggests.
- ・Shipped under MIT, so fine-tuning, resale, and self-hosting need no permission or attribution.
Weaknesses
- ・Self-hosting is not trivial: a practical 4-bit build still needs ~190GB of combined RAM/VRAM, not a single consumer GPU.
- ・The 100k-domestic-chip cluster claim behind the stealth Ox-Alpha run is marketing-grade — Zhipu named no hardware and it is unverified.
- ・Capabilities still trail the top closed flagships on the hardest reasoning and agentic tests; treat 57 AA Index as frontier-adjacent, not frontier-leading.
- ・Steep launch promo (input at $0.075) expires Sep 9, after which the real $0.15/$0.50 price is the baseline.
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| GLM-5.3 Flash | 57 (AA Index) | 63.4 (Z.ai Code Bench) | N/A | $0.15/$0.50 |
| Claude Opus 4.8 | 57 (AA Index) | 58.0 (Z.ai Code Bench) | N/A | $10/$50 |
| DeepSeek V4 Pro | 53 (AA Index) | N/A | N/A | $1.32/$3.96 |
| Qwen3.8-Max-0902 | #1 WebDev | 70.0 (QwenSWE V2) | N/A | $2/$6 |
GLM-5.3 Flash, released August 26, 2026 by Zhipu AI (Z.ai), is a 320B-parameter MoE (18B active) open-weight model that debuted under the anonymous Ox-Alpha handle and promptly topped OpenRouter's weekly charts. Priced at $0.15/$0.50 per million tokens — about 1/40 of Claude Opus 4.8 for near-equal coding quality — it is the clearest signal yet that China's open-weight labs are competing on price-performance, not parameter bragging rights. The model is the first native-multimodal GLM-5 and uses a sparse + linear attention hybrid to make 1M-context serving cheap.
Sources
Analysis generated: 2026-09-06