アリババ독점

Qwen3.7-Plus-Preview

이 모델 비교

Alibaba가 개발한 최신 에이전트 특화 모델. 고도의 에이전트 워크플로우와 추론 능력을 구현합니다.

파라미터

Undisclosed

컨텍스트

256K

라이선스

Proprietary

출시일

2026-05-20

일본어 처리 능력

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API 가격

이 모델의 API 가격 정보는 현재 공개되지 않았습니다

강점

    약점

      활용 사례

        심층 분석

        Arena Elo

        1460

        #15 overall, #12 in coding (BenchLM provisional leaderboard)

        SWE-Bench Verified

        77.7%

        vs Qwen3.7-Max: 80.4%

        ScreenSpot Pro (GUI grounding)

        79.0

        Frontier-tier for computer use agents

        Input Price

        $0.40/1M tokens

        ~6x cheaper than Qwen3.7-Max

        Context Window

        1M tokens

        Shared across text, image, and video

        Terminal-Bench 2.0

        70.3%

        Slightly ahead of Qwen3.7-Max (69.7%)

        강점

        • Exceptional GUI grounding (ScreenSpot Pro 79.0) for computer use and browser automation
        • Multimodal interactive hybrid agent unifying vision, code generation, and GUI/CLI operation
        • Budget-tier pricing with $0.40/1M input tokens while maintaining near-Max text performance

        약점

        • Proprietary API-only model with no open weights for self-hosting or customization
        • Slightly trails text-only Qwen3.7-Max on pure-text and math benchmarks (e.g., SWE-Bench Verified 77.7% vs 80.4%)
        • Visual tokens consume shared 1M context budget, making heavy image/video workloads expensive

        경쟁사 비교

        ModelArenaSWEGPQAPrice
        Qwen3.7-Max147580.4%92.4%$2.50/$7.50
        Claude Sonnet 4.61480~77% (estimated)~90% (estimated)$3/$15
        DeepSeek V4 Pro (Max)1485~80% (estimated)~91% (estimated)$1.50/$4.50

        Qwen3.7-Plus is Alibaba's flagship multimodal agent model, released in June 2026 as the vision-capable counterpart to the text-only Qwen3.7-Max. It builds on the Qwen3.7 backbone to deliver frontier-tier GUI grounding (ScreenSpot Pro 79.0) and hybrid GUI+CLI agent capabilities, enabling end-to-end automation of tasks like browser navigation, UI recreation, and software development from screenshots. The model uniquely combines vision perception, code generation, and tool use within a single agent loop, allowing it to "see, think, write, act, and verify" in closed-loop execution.

        Positioned as a "budget multimodal" tier, Plus is roughly six times cheaper than Qwen3.7-Max on input tokens while matching or slightly trailing it on pure-text benchmarks. Its 1M-token context window and 35-hour autonomous run ceiling support long-horizon agent workflows, and it generalizes across agent frameworks like Claude Code, OpenClaw, and Qwen Code. However, it remains API-only with no open weights, which may limit adoption in regulated or air-gapped environments. The model represents Alibaba's strategic move to compete with Western frontier models in the high-stakes GUI agent and computer-use space, offering a cost-effective alternative for teams building screen-aware automation.

        분석 생성일: 2026-07-17