Alibaba Qwen's 8B vision-language model from the Qwen3-VL series. Processes text, images, and video as input, with 256K context (expandable to 1M). Features 32-language OCR, GUI agent capabilities, and spatial perception. Apache 2.0 license. Released October 15, 2025.
Parameters
8B
Context Window
256K
License
Apache 2.0
Release Date
2025-10-15
API Pricing
Input Price (per 1M tokens)
$0.18
Output Price (per 1M tokens)
$0.7
Billing Mode: per-token
Strengths
Weaknesses
Use Cases
Deep Analysis
Release
October 15, 2025
Qwen3-VL series vision-language model
Parameters
8B
Compact instruction-tuned VLM
Context Window
256K tokens
Long-context vision-language processing
Modalities
Text + Image (+ video) in, text out
Qwen3-VL line
Type
chat / vision-language
Strengths
- ・Compact 8B VLM with a 256K context window - runs on a single consumer GPU while handling long documents and many images.
- ・Instruction-tuned for practical vision tasks: OCR, chart/table understanding, and grounded image Q&A.
- ・Benefits from the Qwen open ecosystem: public weights and repos, easy fine-tuning and self-hosting.
- ・Multimodal input (text, image, and video per the Qwen3-VL line) suits document and UI understanding.
Weaknesses
- ・8B scale means weaker reasoning than flagship VLMs; expect gaps on complex multi-hop visual inference.
- ・No published per-token API pricing in the catalog; cost depends on where you deploy it.
- ・Output is text only - no image or video generation.
- ・Details on the exact benchmark set are scarce; validate on your own vision tasks.
Competitor Comparison
| Model | Arena | GPQA | Price |
|---|---|---|---|
| Qwen3-VL-8B-Instruct (this) | - | - | self-host / varies |
| Qwen3-VL-Plus | - | - | API |
| GLM-4.5V | - | - | open |
Qwen3-VL-8B-Instruct is Alibaba's compact, instruction-tuned vision-language model from the Qwen3-VL series, pairing an 8B parameter body with a 256K-token context for document, chart, and image understanding that fits on a single GPU.
Analysis generated: 2026-09-08