DeepSeekOpen Source

DeepSeek V4.1 Flash

Compare this model

DeepSeek V4.1 Flash is DeepSeek's multimodal flagship refresh, released September 10, 2026. It keeps a 1-million-token context and a new architecture with native image and audio input, decoding at 333–400+ tokens per second. DeepSeek cut cached-input pricing by roughly 60% — as low as ~$0.028 per million tokens in off-peak windows — and now auto-routes DeepSeek V4 Pro traffic to V4.1 Flash. The model ships under an open-weight license and is the centerpiece of DeepSeek's push toward a STAR Market IPO.

Parameters

TBA (new architecture, native multimodal)

Context Window

1M

License

MIT License

Release Date

2026-09-10

API Pricing

Input Price (per 1M tokens)

$0.38

Output Price (per 1M tokens)

$

Billing Mode: standard

Strengths

  • 1M-token context with native multimodal (image/audio) input
  • Decode 333–400+ tok/s, among the fastest open-weight flagships
  • Cached-input price cut ~60%, off-peak as low as ~$0.028/1M tokens
  • Open-weight license; V4 Pro traffic auto-routed to it

Weaknesses

  • Independent benchmark numbers not published at launch
  • Multimodal quality vs top closed models unproven
  • Throughput claims are vendor figures, not third-party runs

Use Cases

  • High-volume agent and coding pipelines needing cheap long context
  • Low-cost multimodal chat and document understanding
  • Self-hosted open-weight deployment with on-prem data

Deep Analysis

Release

Sep 10, 2026

DeepSeek V4.1 Flash; native multimodal refresh

Context Window

1M tokens

Native multimodal (image + audio) input

Decode speed

333–400+ tok/s

Vendor figure; among fastest open-weight flagships

Cached-input price

~$0.028/1M (off-peak)

60% cut vs V4 Pro; ~¥0.02/1M

License

Open weight (MIT)

V4 Pro traffic now auto-routes to it

Intelligence Index

TBA

Independent benchmarks not yet published at launch

Strengths

  • Keeps a 1M-token context with native image and audio input in a new architecture
  • Decodes at 333–400+ tok/s — among the fastest open-weight flagships, good for high-volume agent loops
  • Cached-input price cut ~60%, off-peak as low as ~$0.028/1M tokens; V4 Pro traffic auto-routes here
  • Open-weight license keeps it self-hostable and cheap to run at scale

Weaknesses

  • Independent benchmark numbers were not published at launch — verify on your own tasks
  • Multimodal quality versus top closed models (GPT-6 Astra, Gemini) is unproven
  • Throughput claims are vendor figures, not third-party runs

Competitor Comparison

ModelArenaSWEGPQAPrice
DeepSeek V4 Flash-SWE-bench Verified ~79%GPQA ~91%~$0.094 blended
GPT-5.6 Luna--91.6%$0.450 blended
Qwen3.7 Plus--~88–90%$0.40/$1.16

DeepSeek V4.1 Flash is DeepSeek's multimodal flagship refresh, released Sep 10, 2026. It carries a 1M-token context and a new architecture with native image and audio input, decoding at 333–400+ tokens per second. DeepSeek cut cached-input pricing by roughly 60% (off-peak as low as ~$0.028/1M tokens) and now auto-routes V4 Pro traffic to V4.1 Flash. Shipped under an open-weight license, it is the centerpiece of DeepSeek's push toward a STAR Market IPO. Independent benchmark numbers were not published at launch, so treat vendor claims as provisional.

Analysis generated: 2026-09-11