Moonshot AIProprietary

Kimi K2.5 Instant

Compare this model

A high-performance language model developed by Moonshot AI. It is the strongest model from Moonshot, being open-source with a 1T-parameter MoE architecture.

Parameters

Undisclosed

Context Window

License

Proprietary

Release Date

2026-04-01

Japanese Language Capability

🌐Multilingual

General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.

API Pricing

API pricing for this model is not yet available

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        LM Arena Code Elo

        1526

        #10 overall (Instant variant), sibling K2.5 Thinking at #7 (1550)

        SWE-Bench Verified

        76.8%

        vs Claude Opus 4.5: 80.9%, GPT-5.2: 80.0%

        AIME 2025

        96.1%

        Thinking mode; vs GPT-5.2: 100%

        Input Price

        $0.60/1M

        ~25x cheaper than Claude Opus 4.5 ($15/1M)

        Context Window

        256K tokens

        vs Gemini 3 Pro: 1M, Claude 4.5: 200K

        GDPval-AA Elo

        1309

        #2 open-weight model behind frontier labs only

        Total / Active Parameters

        1T / 32B

        MoE: 384 experts, 8 selected per token

        Strengths

        • Industry-leading Agent Swarm capability: up to 100 parallel sub-agents delivering ~4.5x speedups on parallelizable tasks
        • Aggressively low API pricing ($0.60/1M input) with open-weight availability under Modified MIT license
        • Native multimodal architecture (MoonViT) trained end-to-end on 15T vision-text tokens, excelling at image/video understanding and vision-to-code workflows

        Weaknesses

        • Extreme output verbosity (~6x median token generation), inflating real-world costs and increasing latency despite cheap per-token pricing
        • High self-hosting demands: INT4 quantized weights still require ~630GB storage and 8× H100/H200 GPUs; attention layers remain in BF16
        • Tool-call reliability issues: ~12% failure rate in agent workflows, with occasional goal drift and inconsistent routing outputs

        Competitor Comparison

        ModelArenaSWEGPQAPrice
        GPT-5.2 (xhigh)~1580*80.0%92.4%$2.50/$10.00
        Claude Opus 4.5~1570*80.9%87.0%$15.00/$75.00
        Gemini 3 Pro~1560*76.2%91.9%$1.25/$5.00
        DeepSeek V3.2~1540*73.1%82.4%$0.27/$1.10

        Kimi K2.5 Instant is the fast-response variant of Moonshot AI's flagship open-weight model, a 1-trillion-parameter Mixture-of-Experts architecture that activates only 32 billion parameters per token. Released January 27, 2026, it represents the first Moonshot model to offer native multimodal (image and video) inputs alongside text, making it the leading open-weight model with integrated vision capabilities. The Instant variant specifically disables chain-of-thought reasoning to prioritize speed and cost efficiency, targeting interactive applications like chatbots, quick code generation, and real-time assistants.

        The model's most distinctive innovation is Agent Swarm — a capability unique among frontier models — which orchestrates up to 100 parallel sub-agents using Parallel-Agent Reinforcement Learning (PARL). This enables ~4.5× speedups on parallelizable research and analysis tasks. Combined with aggressive API pricing ($0.60 per million input tokens, roughly 25× cheaper than Claude Opus 4.5), Kimi K2.5 Instant occupies a compelling niche for teams needing frontier-class reasoning at open-weight economics.

        However, the model carries real trade-offs. Output generation is extremely verbose (~89 million tokens across Artificial Analysis's evaluation suite, roughly 6× the median model), which partially erodes the cost advantage. Self-hosting requires substantial GPU infrastructure (~630GB for INT4 weights, 8× high-end GPUs), agent workflows suffer ~12% tool-call failure rates, and English writing quality lags behind Claude and GPT. The 256K context window, while large, trails Gemini's 1M window by 4×. For teams willing to manage these constraints — particularly those building agentic or multimodal pipelines — K2.5 Instant offers capabilities unavailable from any competing open-weight model.

        Analysis generated: 2026-07-17