DeepSeek V4.1 Flash is DeepSeek's multimodal flagship refresh, released September 10, 2026. It keeps a 1-million-token context and a new architecture with native image and audio input, decoding at 333–400+ tokens per second. DeepSeek cut cached-input pricing by roughly 60% — as low as ~$0.028 per million tokens in off-peak windows — and now auto-routes DeepSeek V4 Pro traffic to V4.1 Flash. The model ships under an open-weight license and is the centerpiece of DeepSeek's push toward a STAR Market IPO.
Parameters
TBA (new architecture, native multimodal)
Context Window
1M
License
MIT License
Release Date
2026-09-10
API Pricing
Input Price (per 1M tokens)
$0.38
Output Price (per 1M tokens)
$
Billing Mode: standard
Strengths
- •1M-token context with native multimodal (image/audio) input
- •Decode 333–400+ tok/s, among the fastest open-weight flagships
- •Cached-input price cut ~60%, off-peak as low as ~$0.028/1M tokens
- •Open-weight license; V4 Pro traffic auto-routed to it
Weaknesses
- •Independent benchmark numbers not published at launch
- •Multimodal quality vs top closed models unproven
- •Throughput claims are vendor figures, not third-party runs
Use Cases
- •High-volume agent and coding pipelines needing cheap long context
- •Low-cost multimodal chat and document understanding
- •Self-hosted open-weight deployment with on-prem data
Deep Analysis
Release
Sep 10, 2026
DeepSeek V4.1 Flash; native multimodal refresh
Context Window
1M tokens
Native multimodal (image + audio) input
Decode speed
333–400+ tok/s
Vendor figure; among fastest open-weight flagships
Cached-input price
~$0.028/1M (off-peak)
60% cut vs V4 Pro; ~¥0.02/1M
License
Open weight (MIT)
V4 Pro traffic now auto-routes to it
Intelligence Index
TBA
Independent benchmarks not yet published at launch
Strengths
- ・Keeps a 1M-token context with native image and audio input in a new architecture
- ・Decodes at 333–400+ tok/s — among the fastest open-weight flagships, good for high-volume agent loops
- ・Cached-input price cut ~60%, off-peak as low as ~$0.028/1M tokens; V4 Pro traffic auto-routes here
- ・Open-weight license keeps it self-hostable and cheap to run at scale
Weaknesses
- ・Independent benchmark numbers were not published at launch — verify on your own tasks
- ・Multimodal quality versus top closed models (GPT-6 Astra, Gemini) is unproven
- ・Throughput claims are vendor figures, not third-party runs
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| DeepSeek V4 Flash | - | SWE-bench Verified ~79% | GPQA ~91% | ~$0.094 blended |
| GPT-5.6 Luna | - | - | 91.6% | $0.450 blended |
| Qwen3.7 Plus | - | - | ~88–90% | $0.40/$1.16 |
DeepSeek V4.1 Flash is DeepSeek's multimodal flagship refresh, released Sep 10, 2026. It carries a 1M-token context and a new architecture with native image and audio input, decoding at 333–400+ tokens per second. DeepSeek cut cached-input pricing by roughly 60% (off-peak as low as ~$0.028/1M tokens) and now auto-routes V4 Pro traffic to V4.1 Flash. Shipped under an open-weight license, it is the centerpiece of DeepSeek's push toward a STAR Market IPO. Independent benchmark numbers were not published at launch, so treat vendor claims as provisional.
Analysis generated: 2026-09-11