Meta AIプロプライエタリ

Muse Spark by Meta Superintelligence Labs

このモデルを比較

Meta AI開発の推論特化モデル。高度な推論タスクに優れた性能。

シェア:XはてブLINE

パラメータ

非公開

コンテキスト長

ライセンス

プロプライエタリ

リリース日

2026-04-08

ベンチマーク性能

AA インテリジェンス指数

LMArena Elo

Human-Like Evaluation

ARC-AGI-2

SWE-bench Verified

GPQA Diamond

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

日本語性能

🌐多言語対応

一般的な多言語対応モデル。基本的な日本語処理は可能だが、特化モデルには劣る。

API料金

このモデルのAPI料金情報は現在未公開です

強み

    弱み

      活用例

        深度分析

        Intelligence Index (AA v4.0)

        51

        Tied with GPT-5.4, behind Grok 4.5 (54) and Claude Fable 5 (60)

        Humanity's Last Exam (HLE)

        62.1%

        With tools, Muse Spark 1.1 - vs GPT 5.5: 52.2%, Opus 4.8: 57.9%

        HealthBench Hard

        42.8%

        Leads GPT-5.4 (40.1%) and Gemini 3.1 Pro (20.6%)

        Context Window

        1M tokens

        Up from 262k for Muse Spark 1.0

        API Pricing

        $1.25/$4.25 per 1M tokens

        Input/Output, cache hits at $0.15/1M

        Token Efficiency

        94M tokens for AA Index

        vs GPT-5.4 (109M), GLM-5.2 (141M)

        強み

        • Best-in-class health and medical reasoning with physician-curated training data
        • Unique multi-agent 'Contemplating' mode for parallel reasoning on complex problems
        • Exceptional token efficiency and cost-effective inference at scale

        弱み

        • Significantly trails GPT-5.4 and Claude in coding and agentic tasks (ARC-AGI-2: 42.5 vs 76.1)
        • API access limited to private preview for most developers
        • Weak in abstract reasoning and autonomous desktop workflows (GDPval-AA Elo: 1444 vs 1672)

        競合比較

        ModelArenaSWEGPQAPrice
        GPT-5.457 (AA Index)57.7% (SWE-Bench Pro)92.8% (GPQA Diamond)$2.50/$20 per 1M tokens
        Claude Opus 4.653 (AA Index)80.8% (SWE-Bench Verified)92.7% (GPQA Diamond)$5/$25 per 1M tokens
        Gemini 3.1 Pro57 (AA Index)54.2% (SWE-Bench Pro)94.3% (GPQA Diamond)$2/$12 per 1M tokens

        Muse Spark is the debut model from Meta Superintelligence Labs, representing a strategic pivot from Meta's open-source Llama lineage to a proprietary, natively multimodal reasoning architecture. Launched in April 2026, the model was built from the ground up over nine months following Meta's $14.3 billion investment in Scale AI and the establishment of MSL under Alexandr Wang. It scores 51 on the Artificial Analysis Intelligence Index, placing it among the frontier tier but behind GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6 on general benchmarks.

        The model's standout innovations include its multi-agent 'Contemplating' mode—which orchestrates parallel reasoning agents to achieve superior performance on complex tasks like Humanity's Last Exam (62.1% with tools)—and exceptional performance in health and medical AI, where it leads all competitors with a 42.8% score on HealthBench Hard. Architecturally, Muse Spark is natively multimodal from the ground up, integrating text, image, and audio processing rather than bolting vision onto a language backbone. This enables unique capabilities like visual chain-of-thought reasoning and interactive health visualizations.

        While Muse Spark excels in health, scientific reasoning, and multimodal perception, it currently lags significantly in coding, abstract reasoning, and agentic computer-use tasks compared to GPT-5.4 and Claude. The model is available for free to consumers through Meta's apps, with a private API preview for developers. Meta positions it as the first step on a scaling ladder toward 'personal superintelligence,' with larger models already in development. Its token efficiency—using roughly half the output tokens of competitors for comparable benchmark runs—makes it economically viable for Meta to offer at no cost across its 3+ billion user base.

        分析生成日: 2026-07-17