A foundation model developed by MistralAI.
Parameters
6750
Context Window
256K
License
Apache 2.0
Release Date
2025-12-02
Japanese Language Capability
Multilingual model with strong Japanese language processing capabilities.
API Pricing
Input Price (per 1M tokens)
$0.5
Output Price (per 1M tokens)
$
Billing Mode: standard
Strengths
- •Latest version of Mistral's large-scale model series.
- •Known for high performance across various benchmarks.
- •Offers both base and instruction-tuned variants.
Weaknesses
- •Requires significant computational resources due to its large size.
- •Faces competition from other leading large language models.
- •Specific pricing and access policies may vary.
Use Cases
- •High-performance natural language understanding and generation.
- •Complex instruction following and task execution.
- •Research and enterprise applications requiring top-tier models.
Deep Analysis
Arena Elo
1418
#2 open-source non-reasoning, #6 OSS overall
MMLU
85.5%
Multilingual knowledge benchmark
HumanEval
92.0%
Python code generation
Context Window
256K tokens
Largest among frontier-class open models
Input Price
$0.50/1M tokens
~6x cheaper than Claude Sonnet
Parameters
41B active / 675B total
Granular MoE architecture
Strengths
- ・Apache 2.0 license enables unrestricted commercial use and self-hosting
- ・Best-in-class multilingual support (40+ languages) for European/enterprise workloads
- ・Large 256K context window with native multimodal (text + images) capabilities
Weaknesses
- ・Not a dedicated reasoning model; lags behind chain-of-thought models on complex reasoning (GPQA Diamond 43.9%)
- ・Behind vision-first models on multimodal tasks despite integrated vision encoder
- ・Complex deployment requires substantial infrastructure (8xH200 or similar hardware)
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| DeepSeek V3.2 | 1456 | N/A | 82.4% | $0.28/$0.42 |
| Claude Opus 4.5 | 1496 | N/A | 91.3% | $5.00/$25.00 |
| Llama 3.1 405B | 1352 | N/A | 50.7% | $0.89/$0.89 |
Mistral Large 3 represents Mistral AI's most capable model to date, combining a 675B total parameter granular Mixture-of-Experts architecture with 41B active parameters per token. Released in December 2025 under the Apache 2.0 license, it positions itself as a frontier-class open-weight model optimized for enterprise multilingual pipelines, structured output reliability, and cost-efficient inference. With a 256K context window, native multimodal capabilities (text and images), and strong function calling, it targets production-grade applications like long document understanding, RAG systems, and multilingual customer support.
The model's strategic positioning emphasizes European data sovereignty, with training on 3,000 NVIDIA H200 GPUs and deployment partnerships with major cloud providers. While it achieves strong benchmark scores on general knowledge (MMLU 85.5%) and code generation (HumanEval 92%), it explicitly lacks chain-of-thought reasoning, placing it behind specialized reasoning models on PhD-level scientific questions (GPQA Diamond 43.9%). Its value proposition combines frontier capabilities at $0.50/$1.50 per million tokens—approximately 10x cheaper than Claude Opus 4.6—making it compelling for high-volume enterprise workloads where reasoning depth is secondary to reliability, multilingual support, and cost efficiency.
Sources
- Mistral Large 3 675B Instruct 2512 - Hugging Face Model Card
- Introducing Mistral 3 | Mistral AI Blog
- Model Selection Guide | Mistral Docs
- Mistral Large 3 | Awesome Agents
- Mistral Large 3 675B Instruct – 128k context, multimodal, open source | LLM Reference
- Mistral Large 3 2512 - Pricing & Benchmarks 2026 | LM Market Cap
- Mistral Large 3 — 675B MoE, Apache 2.0, and the Value Play for Enterprise — ChatForest
- Mistral Large 3 vs GPT-5.1 vs Claude vs Gemini: Benchmark (2026) | AI Crucible
- Llama 3.1 405B Instruct vs Mistral Large 3 (675B Instruct 2512) | LLM Stats
- Mistral Large 3: 675B Parameter MoE Model Released Open Source | TPS
Analysis generated: 2026-07-17