A high-performance foundation model developed by Zhipu AI. It excels in Chinese language support and handles a variety of tasks.
Parameters
Undisclosed
Context Window
200K
License
Proprietary
Release Date
2026-04-02
Benchmark Performance
AA Intelligence Index
—
LMArena Elo
—
HLE
—
ARC-AGI-2
—
SWE-bench Verified
—
GPQA Diamond
—
MMLU-Pro
—
LiveCodeBench
—
AIME 2025
—
MATH-500
—
Japanese Language Capability
General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.
API Pricing
API pricing for this model is not yet available
Strengths
Weaknesses
Use Cases
Deep Analysis
Design2Code Score
94.8
vs Claude Opus 4.6: 77.3
Input Price (OpenRouter)
$1.20/1M tokens
Cheapest tracked provider route
Output Price (OpenRouter)
$4.00/1M tokens
Cache read: $0.24/1M
Context Window
200K tokens
Max output: 131K tokens
Architecture
MoE
744B total, 40B active parameters
Release Date
April 1, 2026
Z.AI's first native multimodal coding model
Strengths
- ・Native multimodal perception integrated as core reasoning component, not auxiliary interface
- ・Strong performance on multimodal coding benchmarks (94.8 on Design2Code) while preserving text-only coding capability
- ・Broad joint RL optimization across 30+ task categories yields robust gains in perception, reasoning, and agentic execution
Weaknesses
- ・Slower inference speed compared to text-only models (34th percentile for output speed per benchable.ai)
- ・Pricing varies significantly across providers ($0.70 to $5.00/1M input tokens)
- ・Limited independent benchmark validation; key evaluations like ZClawBench are Z.AI's proprietary benchmarks
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| Claude Opus 4.6 | N/A | N/A | N/A | $3.00/$15.00 |
| Kimi K2.5 | N/A | N/A | N/A | N/A |
| GPT-5.2 | N/A | 80.0%* | N/A | N/A |
GLM-5V-Turbo is Z.AI's first native multimodal agent foundation model, designed specifically for vision-based coding and agent-driven tasks. Unlike models that bolt vision capabilities onto language backends, GLM-5V-Turbo integrates multimodal perception as a core component of reasoning, planning, tool use, and execution. The model represents a significant architectural advancement with innovations like the CogViT vision encoder and Multimodal Multi-Token Prediction (MMTP) that enable efficient processing of images, video, and text while maintaining training and inference efficiency.
Positioned at the intersection of multimodal understanding and agentic execution, GLM-5V-Turbo excels at tasks requiring long-horizon planning, complex coding, and real-world interaction. It demonstrates strong performance on multimodal coding benchmarks while preserving competitive text-only coding capability, making it versatile for both visual and text-first workflows. The model is optimized for integration with agent frameworks like Claude Code and OpenClaw, enabling complete perception–planning–execution loops for practical applications.
Key innovations include hierarchical optimization strategies that build agent capabilities across multiple levels of complexity, and an expanded multimodal toolchain that supports functions from image recognition to deep research. The development process also yielded important insights for agentic model development, emphasizing the foundational role of perception and the efficiency gains from hierarchical optimization over monolithic training approaches.
Sources
- GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
- GLM-5V-Turbo - Overview - Z.AI DEVELOPER DOCUMENT
- GLM 5V Turbo - API Pricing & Benchmarks | OpenRouter
- Z.ai: GLM 5V Turbo - AI Model Details & Benchmarks
- GLM-5V-Turbo – 200k context, multimodal, open source | LLM Reference
- GLM-5-Turbo vs GLM-5V-Turbo: Which Agent Model to Use - Verdent Guides
- GLM-5V-Turbo: Pricing, Context (200K) & Providers · AI Model Intelligence
- Claude Sonnet 4.6 vs GLM-5V-Turbo Comparison (2026) | LLMReference
Analysis generated: 2026-07-17