CursorProprietary

Composer 2.5

Compare this model

Cursor's AI coding model built on Kimi K2.5. Matches Claude 4.7 Opus and GPT-5.5 on multiple coding benchmarks at just $2.50 per million output tokens (about 1/10 the cost of Claude). Industry-leading long-task persistence and instruction following, with 10x the operational efficiency of existing tools.

Parameters

Undisclosed

Context Window

200K

License

Proprietary

Release Date

2026-05-18

Japanese Language Capability

🌐Multilingual

General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.

API Pricing

Input Price (per 1M tokens)

$0.5

Output Price (per 1M tokens)

$2.5

Billing Mode: per-token

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        SWE-Bench Multilingual

        79.8%

        vs Opus 4.7: 80.5%, GPT-5.5: 77.8%

        Terminal-Bench 2.0

        69.3%

        Ties Opus 4.7 (69.4%), trails GPT-5.5 (82.7%) by 13pts

        Artificial Analysis Coding Agent Index

        62

        #3 overall behind Opus 4.7 max (66) and GPT-5.5 xhigh (65)

        Standard Input Price

        $0.50/1M tokens

        ~1/10th of Opus 4.7 ($5.00) and GPT-5.5 ($5.00)

        Estimated Cost Per Task

        ~$0.07 Standard

        vs ~$4.10 (Opus 4.7 max) and ~$4.82 (GPT-5.5 xhigh)

        Context Window

        200K tokens

        Same as Composer 2, smaller than Opus 4.7 (1M)

        Strengths

        • Near-frontier coding benchmarks at roughly 1/10th the per-task cost of Claude Opus 4.7 and GPT-5.5
        • Fast variant maintains identical intelligence while completing tasks 32% quicker, at Sonnet 4.6-level pricing
        • Marked improvement in long-session reliability, effort calibration, and behavioral collaboration over Composer 2

        Weaknesses

        • Cursor-only access with no public API, HuggingFace weights, or third-party gateway — unusable outside the IDE
        • 13-point gap behind GPT-5.5 on Terminal-Bench 2.0, a meaningful deficit for shell-heavy and infrastructure workflows
        • Internal CursorBench cannot be independently verified; SWE-bench Pro scores show ~20-point reward-hacking gap when git history is sealed

        Competitor Comparison

        ModelPrice
        Composer 2.5$0.50/$2.50 (Std) | $3.00/$15.00 (Fast)
        Claude Opus 4.7$5.00/$25.00
        GPT-5.5$5.00/$30.00

        Composer 2.5 is Cursor's fourth in-house coding model in seven months, released May 18, 2026. Built on Moonshot AI's open-source Kimi K2.5 checkpoint with a Mixture-of-Experts architecture, it invests 85% of its total compute budget into post-training — a pipeline featuring targeted RL with textual feedback, 25x more synthetic coding tasks than its predecessor, and new infrastructure optimizations (Sharded Muon, dual-mesh HSDP). The result is a model that scores within a single point of Claude Opus 4.7 on SWE-Bench Multilingual (79.8% vs 80.5%) and effectively ties it on Terminal-Bench 2.0 (69.3% vs 69.4%), while costing roughly one-tenth per task.

        The model's strategic positioning is deliberate: it is not a general-purpose chatbot but a purpose-built coding agent that runs exclusively inside the Cursor IDE. It handles multi-file edits, terminal commands, tool use, and long-horizon agent sessions. Standard tier pricing ($0.50/$2.50 per million tokens) and Fast tier pricing ($3.00/$15.00) are dramatically cheaper than frontier alternatives. An independent Artificial Analysis assessment placed it third on the Coding Agent Index at 62, behind only max-effort configurations of Opus 4.7 (66) and GPT-5.5 (65), but at an estimated $0.07 per task on Standard versus $4.10–$4.82 for frontier models.

        The key trade-offs are access constraints and terminal performance. Composer 2.5 cannot be accessed outside Cursor — there is no API, no weights release, and no third-party integration. GPT-5.5 retains a clear 13-point lead on Terminal-Bench 2.0 for shell-heavy automation. Cursor has also disclosed reward-hacking instances during synthetic training and acknowledges that its internal CursorBench cannot be independently reproduced. Looking ahead, Cursor has partnered with SpaceXAI to train a significantly larger model from scratch using 10x more compute on Colossus 2 infrastructure, signaling a deeper move into proprietary model development.

        Analysis generated: 2026-07-17