MiniMaxAI오픈소스

MiniMax M2.5

이 모델 비교

MiniMax가 개발한 고성능 MoE 모델. 230B 파라미터이며, 100만 토큰 컨텍스트를 지원합니다.

파라미터

2290

컨텍스트

128K

라이선스

https://github.com/MiniMax-AI/MiniMax-M2.5/blob/main/LICENSE-MODEL

출시일

2026-02-12

벤치마크 성능

AA Intelligence Index

LMArena Elo

HLE

ARC-AGI-2

SWE-bench Verified

GPQA Diamond

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

일본어 처리 능력

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API 가격

입력 가격 (1M 토큰당)

$0.3

출력 가격 (1M 토큰당)

$

과금 모드: standard

강점

    약점

      활용 사례

        심층 분석

        SWE-Bench Verified

        80.2%

        Highest among open-weight models, 0.6 points behind Claude Opus 4.6

        Multi-SWE-Bench

        51.3%

        First place, ahead of Claude Opus 4.6's 50.3%

        BrowseComp

        76.3%

        Leading open-weight model for web search tasks with context management

        Input Price (Standard)

        $0.15/1M tokens

        10-20x cheaper than Claude Opus 4.6

        Output Price (Standard)

        $1.20/1M tokens

        Drastically lower cost for high-volume agentic tasks

        Speed (Lightning)

        100 tokens/sec

        Nearly twice that of other frontier models like Claude Opus 4.6 (~60 tokens/sec)

        강점

        • Exceptional cost-performance ratio for coding and agentic tasks, with SOTA benchmarks at frontier prices
        • Fast inference up to 100 tokens/sec for real-time agentic workflows and low latency
        • Open weights under MIT license enabling self-hosting, fine-tuning, and data sovereignty

        약점

        • High hallucination rate of 88% on AA-Omniscience evaluation, reducing factual reliability
        • Lags in creative reasoning and hard math benchmarks (e.g., AIME 2025: 86.3 vs 95.6+ for frontier models)
        • Text-only model without multimodal support, limiting use cases requiring image or video processing

        경쟁사 비교

        ModelArenaSWEGPQAPrice
        Claude Opus 4.6~53 (Intelligence Index)80.8%90.0%$5.00/$25.00 per 1M tokens
        GPT-5.2~57 (Intelligence Index)80.0%90.0%~$1.25/$10.00 per 1M tokens
        Gemini 3 Pro~57 (Intelligence Index)78%91.0%$2.00/$8.00 per 1M tokens

        MiniMax M2.5 is a high-performance Mixture-of-Experts model with 230 billion total parameters and 10 billion active, released on February 12, 2026. It achieves state-of-the-art results in coding, agentic tool use, and office productivity, scoring 80.2% on SWE-Bench Verified and 76.3% on BrowseComp, while being dramatically cost-effective at $0.15 per million input tokens for the standard variant. Trained via large-scale reinforcement learning across 200,000+ real-world environments using MiniMax's Forge framework, M2.5 excels in task decomposition and efficiency, completing agentic tasks 37% faster than its predecessor M2.1.

        Positioned as a frontier-competitive open-weight model, M2.5 targets developers and enterprises needing high-volume, cost-sensitive coding and automation workflows. Its MIT license allows for self-hosting and customization, challenging closed-source models by offering near-matching performance at a fraction of the cost. However, it has notable limitations including a high hallucination rate and lack of multimodal capabilities, making it a specialist for coding and agentic tasks rather than a general-purpose model.

        분석 생성일: 2026-07-17