Meta AI독점

Muse Spark by Meta Superintelligence Labs

이 모델 비교

Meta AI에서 개발한 추론 특화 모델. 고급 추론 작업에서 뛰어난 성능을 발휘합니다.

파라미터

Undisclosed

컨텍스트

라이선스

Proprietary

출시일

2026-04-08

벤치마크 성능

AA Intelligence Index

LMArena Elo

HLE

ARC-AGI-2

SWE-bench Verified

GPQA Diamond

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

일본어 처리 능력

🌐Multilingual

General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.

API 가격

이 모델의 API 가격 정보는 현재 공개되지 않았습니다

강점

    약점

      활용 사례

        심층 분석

        Intelligence Index (AA v4.0)

        51

        Tied with GPT-5.4, behind Grok 4.5 (54) and Claude Fable 5 (60)

        Humanity's Last Exam (HLE)

        62.1%

        With tools, Muse Spark 1.1 - vs GPT 5.5: 52.2%, Opus 4.8: 57.9%

        HealthBench Hard

        42.8%

        Leads GPT-5.4 (40.1%) and Gemini 3.1 Pro (20.6%)

        Context Window

        1M tokens

        Up from 262k for Muse Spark 1.0

        API Pricing

        $1.25/$4.25 per 1M tokens

        Input/Output, cache hits at $0.15/1M

        Token Efficiency

        94M tokens for AA Index

        vs GPT-5.4 (109M), GLM-5.2 (141M)

        강점

        • Best-in-class health and medical reasoning with physician-curated training data
        • Unique multi-agent 'Contemplating' mode for parallel reasoning on complex problems
        • Exceptional token efficiency and cost-effective inference at scale

        약점

        • Significantly trails GPT-5.4 and Claude in coding and agentic tasks (ARC-AGI-2: 42.5 vs 76.1)
        • API access limited to private preview for most developers
        • Weak in abstract reasoning and autonomous desktop workflows (GDPval-AA Elo: 1444 vs 1672)

        경쟁사 비교

        ModelArenaSWEGPQAPrice
        GPT-5.457 (AA Index)57.7% (SWE-Bench Pro)92.8% (GPQA Diamond)$2.50/$20 per 1M tokens
        Claude Opus 4.653 (AA Index)80.8% (SWE-Bench Verified)92.7% (GPQA Diamond)$5/$25 per 1M tokens
        Gemini 3.1 Pro57 (AA Index)54.2% (SWE-Bench Pro)94.3% (GPQA Diamond)$2/$12 per 1M tokens

        Muse Spark is the debut model from Meta Superintelligence Labs, representing a strategic pivot from Meta's open-source Llama lineage to a proprietary, natively multimodal reasoning architecture. Launched in April 2026, the model was built from the ground up over nine months following Meta's $14.3 billion investment in Scale AI and the establishment of MSL under Alexandr Wang. It scores 51 on the Artificial Analysis Intelligence Index, placing it among the frontier tier but behind GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6 on general benchmarks.

        The model's standout innovations include its multi-agent 'Contemplating' mode—which orchestrates parallel reasoning agents to achieve superior performance on complex tasks like Humanity's Last Exam (62.1% with tools)—and exceptional performance in health and medical AI, where it leads all competitors with a 42.8% score on HealthBench Hard. Architecturally, Muse Spark is natively multimodal from the ground up, integrating text, image, and audio processing rather than bolting vision onto a language backbone. This enables unique capabilities like visual chain-of-thought reasoning and interactive health visualizations.

        While Muse Spark excels in health, scientific reasoning, and multimodal perception, it currently lags significantly in coding, abstract reasoning, and agentic computer-use tasks compared to GPT-5.4 and Claude. The model is available for free to consumers through Meta's apps, with a private API preview for developers. Meta positions it as the first step on a scaling ladder toward 'personal superintelligence,' with larger models already in development. Its token efficiency—using roughly half the output tokens of competitors for comparable benchmark runs—makes it economically viable for Meta to offer at no cost across its 3+ billion user base.

        분석 생성일: 2026-07-17