MiniMaxAIProprietary

MiniMax M3

Compare this model

The latest high-performance MoE model developed by MiniMax. It has 428B parameters (23B active), supports a 1 million token context, uses the MSA architecture, and has native multimodal capabilities.

Parameters

Undisclosed

Context Window

License

https://github.com/MiniMax-AI/MiniMax-M2.7/blob/main/LICENSE

Release Date

2026-06-01

Japanese Language Capability

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API Pricing

Input Price (per 1M tokens)

$2.1

Output Price (per 1M tokens)

$

Billing Mode: standard

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        Artificial Analysis Intelligence Index

        55

        Level with Kimi K2.6 (54) and MiMo-V2.5-Pro (54); top open-weights model pending weight release

        SWE-Bench Verified

        80.5%

        8th of 49 models on BenchLM leaderboard

        SWE-Bench Pro

        59.0%

        vs Opus 4.7: 64.3%, GPT-5.5: 58.6%

        GPQA Diamond

        93%

        Up from M2.7's 87% (+6 points)

        Context Window

        1M tokens

        MSA architecture; 9× prefill, 15× decode speedup vs M2

        Input/Output Price

        $0.30/$1.20 per 1M

        ≤512K context; $0.60/$2.40 for 512K–1M

        Strengths

        • First open-weights model combining frontier coding, 1M context, and native multimodality in a single package
        • Dramatically cheaper than closed frontier models (~1/10th to 1/20th the cost of Claude Opus or GPT-5.5)
        • MiniMax Sparse Attention (MSA) makes 1M-token context computationally practical with 9×–15× speedups over full attention

        Weaknesses

        • Inconsistent quality on some tasks—test-writing and reasoning results vary more than Claude Opus 4.8 or GPT-5.5
        • Open weights not yet released at launch; commercial license terms unknown and likely restrictive based on M2.7 precedent
        • Weak abstract/fluid reasoning (ARC-AGI-2 single-digit scores); heavy abstention on AA-Omniscience (only 30.9% of questions attempted)

        Competitor Comparison

        ModelGPQAPrice
        Claude Opus 4.7~95%~$15/$75 per 1M
        GPT-5.5~93%~$10/$30 per 1M
        Gemini 3.1 Pro~90%~$1.25/$5 per 1M
        Kimi K2.6~88%similar range

        MiniMax M3 is a 428B-parameter mixture-of-experts model (23B active) released June 1, 2026, marking a significant leap for MiniMax's M-series. Its headline innovation is MiniMax Sparse Attention (MSA), a two-stage attention mechanism that pre-filters key-value blocks before running full attention on selected regions, enabling practical 1M-token context at 1/20th the per-token compute of its predecessor. M3 is positioned as the first open-weights model to simultaneously offer frontier-tier coding performance, million-token context, and native multimodal input (text, image, video).

        On benchmarks, M3 sits in a competitive middle ground: it slightly edges GPT-5.5 on SWE-Bench Pro (59.0% vs 58.6%), achieves 93% on GPQA Diamond, and scores 83.5% on BrowseComp (surpassing Claude Opus 4.7's 79.3%). It reaches 80.5% on SWE-Bench Verified and 55 on the Artificial Analysis Intelligence Index, placing it level with other leading open-weights models like Kimi K2.6. However, it trails Claude Opus 4.7 on most system-level and agentic benchmarks, and shows weaker performance on abstract reasoning tasks.

        The commercial strategy is aggressive pricing: API costs are $0.30/$1.20 per million input/output tokens (standard, ≤512K context), roughly 1/10th to 1/20th of closed frontier models. Subscription Token Plans range from $20–$120/month with generous token quotas. The model is available via MiniMax API, OpenRouter, and planned self-hosting once weights release on Hugging Face (promised within ~10 days of launch). For developers, M3 represents a new cost-performance tradeoff: not the absolute best model, but potentially the best value for coding, long-context, and agentic workloads.

        Analysis generated: 2026-07-17