Anthropic독점

Claude Opus 4.7

이 모델 비교

Anthropic에서 개발한 고성능 기반 모델. 다양한 작업에 대응하며 균형 잡힌 성능을 제공합니다.

파라미터

Undisclosed

컨텍스트

라이선스

Proprietary

출시일

2026-04-16

벤치마크 성능

AA Intelligence Index

LMArena Elo

HLE

ARC-AGI-2

SWE-bench Verified

GPQA Diamond

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

일본어 처리 능력

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API 가격

입력 가격 (1M 토큰당)

$2.5

출력 가격 (1M 토큰당)

$

과금 모드: standard

강점

    약점

      활용 사례

        심층 분석

        LMArena Elo

        1492

        WebDev Elo: 1560; Vision Elo: 1304

        SWE-bench Verified

        87.6%

        Up from 80.8% on Opus 4.6

        SWE-bench Pro

        64.3%

        vs GPT-5.4: 57.7%

        GPQA Diamond

        94.2%

        Near parity with GPT-5.4 (94.4%)

        Input/Output Price

        $5 / $25 per 1M tokens

        Same rate card as Opus 4.6; new tokenizer inflates 10–35%

        Context Window

        1M input / 128K output

        Max image resolution: 3.75MP (2,576px long edge)

        강점

        • Industry-leading agentic coding: 87.6% SWE-bench Verified, 64.3% SWE-bench Pro — best-in-class among broadly available models
        • 3.3× higher image resolution (3.75MP) unlocks dense screenshot analysis, diagram extraction, and computer-use agents
        • Strongest multi-tool orchestration at 77.3% MCP-Atlas (+9 points over GPT-5.4), ideal for complex agent pipelines
        • Self-verification before reporting and literal instruction following improve reliability on long-running autonomous tasks

        약점

        • Severe long-context retrieval regression: MRCR v2 8-needle at 1M tokens drops from 78.3% (Opus 4.6) to 32.2%
        • New tokenizer inflates real-world costs 10–35% despite unchanged per-token rate card
        • BrowseComp regressed to 79.3% from 84.0%; trails GPT-5.4 (89.3%) and Gemini 3.1 Pro (85.9%) on web research tasks

        경쟁사 비교

        ModelArenaSWEGPQAPrice
        Claude Opus 4.7149287.6%94.2%$5/$25 per 1M tokens
        GPT-5.4N/A from sources84.1%94.4%Not reported in sources
        Gemini 3.1 ProN/A from sources80.6%94.3%Not reported in sources

        Claude Opus 4.7, released April 16, 2026, is Anthropic's most capable broadly available model and a direct upgrade to Opus 4.6. It targets agentic software engineering, high-fidelity vision tasks, and long-running autonomous workflows. The model introduces self-verification behavior (it checks its own outputs before reporting back), literal instruction following, higher-resolution image processing (3.75MP, up from 1.15MP), and improved file-system memory for multi-session agent work. Pricing is unchanged at $5/$25 per million input/output tokens, though a new tokenizer inflates effective costs by 10–35% depending on content type.

        Opus 4.7 sets new benchmarks for Anthropic on SWE-bench Verified (87.6%), SWE-bench Pro (64.3%), and MCP-Atlas (77.3%), with particularly strong gains on the hardest coding problems. Partner testimonials from Cursor, Replit, Notion, Vercel, and others confirm real-world production improvements of 10–15% in task success rates and 3× more production task resolution on some benchmarks. The model introduces a new 'xhigh' effort level between 'high' and 'max', and ships with task budgets (public beta) for cost control on long-running agent loops.

        However, the release comes with notable trade-offs. Long-context multi-needle retrieval collapsed on MRCR v2 (78.3% → 32.2% at 1M tokens), BrowseComp regressed nearly 5 points, and the new tokenizer raises actual costs despite unchanged pricing. Anthropic positions Opus 4.7 below its unreleased Claude Mythos Preview, and describes it as the first broadly released model carrying cybersecurity safeguards from Project Glasswing — including automated blocking of high-risk cyber requests. The system card rates it 'largely well-aligned and trustworthy, though not fully ideal.'

        분석 생성일: 2026-07-17