DeepSeek오픈소스

DeepSeek V4 Pro Base

이 모델 비교

DeepSeek 개발의 고성능 MoE 기초 모델. 1.6T 파라미터(49B 활성화)로 다양한 태스크를 처리합니다.

파라미터

Undisclosed

컨텍스트

라이선스

MIT

출시일

2026-04-24

벤치마크 성능

AA Intelligence Index

LMArena Elo

HLE

ARC-AGI-2

SWE-bench Verified

GPQA Diamond

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

일본어 처리 능력

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API 가격

이 모델의 API 가격 정보는 현재 공개되지 않았습니다

강점

    약점

      활용 사례

        심층 분석

        Total Parameters

        1.6T

        49B activated per token (MoE)

        Context Window

        1M tokens

        Default, no premium tier

        MMLU (5-shot)

        90.1

        vs V3.2-Base: 87.8

        SimpleQA Verified

        55.2%

        25-shot, vs V3.2: 28.3%

        LongBench-V2

        51.5

        1-shot, vs V3.2: 40.2

        Input Price

        $1.74/1M

        via DeepInfra; $0.145 cached

        강점

        • Massive 1.6T-parameter MoE with only 49B active — frontier knowledge density at efficient inference cost
        • Dramatic long-context efficiency: 27% FLOPs and 10% KV cache of V3.2 at 1M tokens via hybrid CSA/HCA attention
        • MIT-licensed open weights enabling unrestricted fine-tuning, self-hosting, and commercial deployment

        약점

        • Base model only — requires post-training (SFT/RL) for instruction-following, chat, or agentic use cases
        • Preview release status with limited BenchLM benchmark coverage (24 of 249 benchmarks); not yet fully evaluated
        • Massive weight footprint (~862 GB FP4+FP8) demands enterprise-grade multi-GPU infrastructure for self-hosting

        경쟁사 비교

        ModelArenaSWEGPQA
        DeepSeek-V3.2-BaseN/AN/A (base)N/A (base)
        DeepSeek-V4-Flash-BaseN/AN/A (base)N/A (base)
        DeepSeek-V4-Pro-BaseN/AN/A (base)N/A (base)

        DeepSeek-V4-Pro-Base is the pre-trained foundation checkpoint of DeepSeek's flagship V4 series, released April 24, 2026 under the MIT license. As a 1.6-trillion-parameter Mixture-of-Experts model with 49 billion parameters activated per token, it represents the largest open-weight base model available and sets new performance ceilings across knowledge, reasoning, coding, and long-context benchmarks among DeepSeek foundation models. The model was pre-trained on over 33 trillion tokens using a novel architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA), achieving dramatic efficiency gains: at 1M-token context, it requires only 27% of the single-token inference FLOPs and 10% of the KV cache compared to its predecessor, DeepSeek-V3.2.

        Key architectural innovations include Manifold-Constrained Hyper-Connections (mHC) for stable deep-layer signal propagation, the Muon optimizer for faster convergence, and FP4+FP8 mixed-precision training for MoE expert weights. The base model demonstrates comprehensive superiority over V3.2-Base across virtually all evaluated domains — achieving 90.1 on MMLU, 73.5 on MMLU-Pro, 76.8 on HumanEval, 92.6 on GSM8K, and 51.5 on LongBench-V2 — while operating with a significantly larger expert pool and supporting native 1M-token contexts.

        As a base model, DeepSeek-V4-Pro-Base is designed for researchers and developers building custom post-training pipelines. It is not intended for direct conversational use; that role falls to the instruct variants (DeepSeek-V4-Pro with Non-think, Think High, and Think Max reasoning modes). The companion DeepSeek-V4-Flash-Base offers a more deployable alternative at 284B total parameters (13B active), achieving most of Pro's quality at roughly one-fifth the parameter count.

        분석 생성일: 2026-07-17