DeepSeekOpen Source

DeepSeek V4 Flash Base

Compare this model

A high-speed MoE foundation model developed by DeepSeek. It has 284B parameters (13B active) and is optimized for rapid processing.

Parameters

Undisclosed

Context Window

License

MIT

Release Date

2026-04-24

Benchmark Performance

AA Intelligence Index

LMArena Elo

HLE

ARC-AGI-2

SWE-bench Verified

GPQA Diamond

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

Japanese Language Capability

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API Pricing

API pricing for this model is not yet available

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        Arena Elo

        Unranked

        Not listed on BenchLM's public leaderboard due to insufficient coverage

        SWE-Bench Verified

        N/A

        Not measured for base model; instruct model achieves 80.6% (Max mode)

        Input Price

        N/A

        Base model not typically served via API; instruct model priced at $0.14/$0.28 per 1M tokens

        Context Window

        1M tokens

        Hybrid CSA/HCA attention enables efficient long-context processing

        Active Parameters

        13B

        From 284B total (Mixture-of-Experts architecture)

        Knowledge Benchmarks

        44.7 avg

        Ranks #68 of 107 models; strongest category but still limited

        Strengths

        • Dramatically improved long-context efficiency with only 10% FLOPs and 7% KV cache vs V3.2 at 1M tokens
        • Open-weight (MIT license) with 1M-token context window for local deployment and fine-tuning
        • Strong coding and math performance in non-thinking mode (90.8% GSM8K, 93.6% CMath)

        Weaknesses

        • Significantly lower knowledge benchmarks than frontier models (46.5% SuperGPQA vs 90%+ for top models)
        • Not yet ranked on most leaderboards due to limited benchmark coverage (24/296 tracked benchmarks)
        • Complex MoE architecture requires substantial hardware (284B parameters, ~150GB for weights alone)

        Competitor Comparison

        ModelArenaSWEGPQAPrice
        DeepSeek V3.2 BaseN/AN/AN/AN/A
        Claude Opus 4.6#880.8%91.3%$5/$25
        GPT-5.4 xHighN/AN/A93.0%~$15/$75

        DeepSeek V4 Flash Base is a high-efficiency Mixture-of-Experts foundation model with 284B total parameters (13B activated) designed specifically for long-context applications. Released as part of the DeepSeek V4 family in April 2026, it represents a significant architectural advancement with its hybrid Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) mechanism, enabling practical 1-million-token context windows with dramatically reduced computational requirements compared to previous models.

        As a base model (pre-trained only, without instruction tuning), DeepSeek V4 Flash Base is positioned as a research artifact and foundation for fine-tuning rather than a direct API-accessible instruct model. While it shows improvements over DeepSeek V3.2 in several knowledge and coding benchmarks, it still significantly trails frontier closed-source models like Claude Opus 4.6 and GPT-5.4 in most categories. The model's primary innovation lies in its efficiency gains for long-context processing, making it theoretically suitable for applications requiring extensive context understanding without prohibitive computational costs.

        The model is open-weight under MIT license, allowing researchers and developers to examine, modify, and deploy it locally. However, practical deployment remains challenging due to the model's size and the specialized hardware requirements for efficient MoE inference. The V4 family introduces several architectural innovations including Manifold-Constrained Hyper-Connections and the Muon optimizer, which contribute to training stability and convergence speed.

        Analysis generated: 2026-07-17