MistralAI오픈소스

Mistral Large 3 (v25.12, Base/Instruct)

이 모델 비교

MistralAI의 파운데이션 모델입니다.

파라미터

6750

컨텍스트

256K

라이선스

Apache 2.0

출시일

2025-12-02

벤치마크 성능

AA Intelligence Index

LMArena Elo

HLE

ARC-AGI-2

SWE-bench Verified

GPQA Diamond

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

일본어 처리 능력

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API 가격

입력 가격 (1M 토큰당)

$0.5

출력 가격 (1M 토큰당)

$

과금 모드: standard

강점

    약점

      활용 사례

        심층 분석

        Arena Elo

        1418

        #2 open-source non-reasoning, #6 OSS overall

        MMLU

        85.5%

        Multilingual knowledge benchmark

        HumanEval

        92.0%

        Python code generation

        Context Window

        256K tokens

        Largest among frontier-class open models

        Input Price

        $0.50/1M tokens

        ~6x cheaper than Claude Sonnet

        Parameters

        41B active / 675B total

        Granular MoE architecture

        강점

        • Apache 2.0 license enables unrestricted commercial use and self-hosting
        • Best-in-class multilingual support (40+ languages) for European/enterprise workloads
        • Large 256K context window with native multimodal (text + images) capabilities

        약점

        • Not a dedicated reasoning model; lags behind chain-of-thought models on complex reasoning (GPQA Diamond 43.9%)
        • Behind vision-first models on multimodal tasks despite integrated vision encoder
        • Complex deployment requires substantial infrastructure (8xH200 or similar hardware)

        경쟁사 비교

        ModelArenaSWEGPQAPrice
        DeepSeek V3.21456N/A82.4%$0.28/$0.42
        Claude Opus 4.51496N/A91.3%$5.00/$25.00
        Llama 3.1 405B1352N/A50.7%$0.89/$0.89

        Mistral Large 3 represents Mistral AI's most capable model to date, combining a 675B total parameter granular Mixture-of-Experts architecture with 41B active parameters per token. Released in December 2025 under the Apache 2.0 license, it positions itself as a frontier-class open-weight model optimized for enterprise multilingual pipelines, structured output reliability, and cost-efficient inference. With a 256K context window, native multimodal capabilities (text and images), and strong function calling, it targets production-grade applications like long document understanding, RAG systems, and multilingual customer support.

        The model's strategic positioning emphasizes European data sovereignty, with training on 3,000 NVIDIA H200 GPUs and deployment partnerships with major cloud providers. While it achieves strong benchmark scores on general knowledge (MMLU 85.5%) and code generation (HumanEval 92%), it explicitly lacks chain-of-thought reasoning, placing it behind specialized reasoning models on PhD-level scientific questions (GPQA Diamond 43.9%). Its value proposition combines frontier capabilities at $0.50/$1.50 per million tokens—approximately 10x cheaper than Claude Opus 4.6—making it compelling for high-volume enterprise workloads where reasoning depth is secondary to reliability, multilingual support, and cost efficiency.

        분석 생성일: 2026-07-17