Zhipu AI (Z.ai)Open Source

GLM-5.3 Flash

Compare this model

GLM-5.3 Flash is Zhipu AI (Z.ai)'s 320B-parameter MoE (18B active) open-weight model, released August 26, 2026. It debuted anonymously as 'Ox-Alpha' and held OpenRouter's weekly top slot before Zhipu confirmed authorship on August 26. A sparse + linear attention hybrid lets it serve a 1.31M-token context cheaply, and it is the first native-multimodal GLM-5 release. Priced at $0.15/$0.50 per million tokens — about 1/40 of Claude Opus 4.8 — it ties Opus 4.8 on the AA Intelligence Index (57) and tops it on the vendor's own Code Bench (63.4). MIT-licensed.

Parameters

320B total / 18B active (MoE, sparse + linear attention hybrid)

Context Window

1.31M (131K max output)

License

MIT

Release Date

2026-08-26

API Pricing

Input Price (per 1M tokens)

$0.15

Output Price (per 1M tokens)

$0.5

Billing Mode: launch promo input $0.075 through 2026-09-09, then standard $0.15/$0.50

Strengths

  • Priced ~1/40 of Claude Opus 4.8 with comparable Artificial Analysis intelligence
  • First native-multimodal GLM-5 (image + video input)
  • Sparse + linear attention hybrid cuts 1M-context serving cost
  • MIT license: fine-tune, resell, or self-host with no constraints

Weaknesses

  • 4-bit self-host still needs ~190GB of combined RAM/VRAM — not a single-GPU model
  • The 100k-domestic-chip serving claim is unverified and reads as marketing
  • AA Index 57 is frontier-adjacent, not frontier-leading on the hardest reasoning and agentic tests
  • Launch promo ($0.075 input) expires Sep 9; the real baseline is $0.15/$0.50

Use Cases

  • High-volume coding, classification, extraction and summarization where cost dominates
  • Document + image multimodal workflows without a second vision model
  • Replacing or downgrading Claude Opus subscriptions at ~1/40 the cost
  • Self-hosted fine-tuning research where the MIT license matters

Deep Analysis

Context Window

1.31M tokens

131K max output; native image + video input; first native-multimodal GLM-5 release

API Price

$0.15 / $0.50 per 1M

Input/output per million tokens; launch promo $0.075 input through Sep 9

AA Intelligence Index

57

Ties Claude Opus 4.8 on Artificial Analysis composite index

Z.ai Code Bench

63.4

Ahead of Claude Opus 4.8 (58.0) and DeepSeek-V4-Vision-Exp (59.3)

License

MIT (open weights)

Weights published on Hugging Face Aug 26, 2026

Parameters

320B total / 18B active

MoE; sparse + linear attention hybrid; first open frontier model with that mix

Strengths

  • Price-performance is the headline: ~1/40 the cost of Claude Opus 4.8 at comparable coding quality on Z.ai Code Bench.
  • First open-weight GLM-5 model with native multimodal input — can judge when to look and steer with visual feedback.
  • Sparse + linear attention hybrid cuts 1M-context serving cost sharply; runs on far less hardware than the parameter count suggests.
  • Shipped under MIT, so fine-tuning, resale, and self-hosting need no permission or attribution.

Weaknesses

  • Self-hosting is not trivial: a practical 4-bit build still needs ~190GB of combined RAM/VRAM, not a single consumer GPU.
  • The 100k-domestic-chip cluster claim behind the stealth Ox-Alpha run is marketing-grade — Zhipu named no hardware and it is unverified.
  • Capabilities still trail the top closed flagships on the hardest reasoning and agentic tests; treat 57 AA Index as frontier-adjacent, not frontier-leading.
  • Steep launch promo (input at $0.075) expires Sep 9, after which the real $0.15/$0.50 price is the baseline.

Competitor Comparison

ModelArenaSWEGPQAPrice
GLM-5.3 Flash57 (AA Index)63.4 (Z.ai Code Bench)N/A$0.15/$0.50
Claude Opus 4.857 (AA Index)58.0 (Z.ai Code Bench)N/A$10/$50
DeepSeek V4 Pro53 (AA Index)N/AN/A$1.32/$3.96
Qwen3.8-Max-0902#1 WebDev70.0 (QwenSWE V2)N/A$2/$6

GLM-5.3 Flash, released August 26, 2026 by Zhipu AI (Z.ai), is a 320B-parameter MoE (18B active) open-weight model that debuted under the anonymous Ox-Alpha handle and promptly topped OpenRouter's weekly charts. Priced at $0.15/$0.50 per million tokens — about 1/40 of Claude Opus 4.8 for near-equal coding quality — it is the clearest signal yet that China's open-weight labs are competing on price-performance, not parameter bragging rights. The model is the first native-multimodal GLM-5 and uses a sparse + linear attention hybrid to make 1M-context serving cheap.

Analysis generated: 2026-09-06