StepFunAI오픈소스

Step 3.7 Flash

이 모델 비교

StepFun AI가 개발한 고속 기초 모델. 낮은 지연 시간과 높은 처리량을 동시에 달성합니다.

파라미터

1980

컨텍스트

256K

라이선스

Apache 2.0

출시일

2026-05-29

일본어 처리 능력

🌐Multilingual

General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.

API 가격

이 모델의 API 가격 정보는 현재 공개되지 않았습니다

강점

    약점

      활용 사례

        심층 분석

        ClawEval-1.1

        67.1

        #1 overall; next best is 59.8 (Claude 4 Opus)

        SWE-Bench PRO

        56.3

        #2 overall behind Claude 4 Opus at 64.3

        SimpleVQA (Search)

        79.2

        #1; beats GPT 5.5 at 79.1

        Input Price

        $0.20/1M tokens

        $0.04/1M on cache hit; ~25× cheaper than Claude Opus

        Throughput

        400 tok/s

        cloud API; ~27 tok/s on DGX Spark locally

        Context Window

        256K tokens

        ~384 pages of text

        강점

        • Best-in-class agent reliability with ClawEval-1.1 score of 67.1, demonstrating superior multi-step tool orchestration and resistance to adversarial traps
        • Advisor Mode delivers 97% of Claude Opus 4.6 coding performance at $0.19/task vs $1.76/task — a genuine 9× cost reduction for production agentic workflows
        • Extensive local deployment flexibility: runs on Mac Studio, DGX Spark, AMD Ryzen AI Max+ 395 with BF16, FP8, NVFP4, and GGUF quantizations available on day one

        약점

        • Significant gap on Terminal-Bench 2.1 (59.5) vs GPT 5.5 (82.7) and GDPVal-AA (45.8) vs frontier models at 63%, limiting complex terminal and general professional task use cases
        • Self-reported benchmark data; many competitor scores are from StepFun's own re-runs rather than standardized third-party evaluations, reducing comparability confidence
        • 128GB minimum hardware for local deployment excludes consumer laptops and most prosumer setups; full BF16 requires 394GB

        경쟁사 비교

        ModelArenaSWEPrice
        Claude Opus 4.6N/A64.3%$5.00/$25.00
        DeepSeek V4 FlashN/A55.6%$0.22/$0.87
        Gemini 3.5 FlashN/A55.1%$0.15/$0.60

        Step 3.7 Flash is a 198B-parameter sparse Mixture-of-Experts vision-language model released by StepFun AI on May 29, 2026. Despite its massive total parameter count, it activates only ~11B parameters per token through expert routing (8 out of 288 experts), delivering near-frontier intelligence at inference costs comparable to small dense models. The model combines a 196B language backbone with a 1.8B native vision encoder, supports a 256K context window, and achieves up to 400 tokens/second throughput on cloud APIs.

        The model's primary differentiation lies in agentic workflow reliability rather than absolute intelligence ceiling. It leads the ClawEval-1.1 benchmark at 67.1 (a 23.5-point jump over its predecessor Step 3.5 Flash) and places second on SWE-Bench PRO at 56.3. Its most innovative feature is Advisor Mode, a two-tier execution strategy where Step 3.7 Flash handles routine tool-calling and iteration while escalating to a larger frontier model only at genuine decision points, achieving approximately 97% of Claude Opus 4.6's coding performance at roughly one-ninth the per-task cost ($0.19 vs $1.76).

        The model is fully open-weight under Apache 2.0 and ships with quantized checkpoints (BF16, FP8, NVFP4, GGUF) across HuggingFace, enabling local deployment on 128GB+ unified memory devices like NVIDIA DGX Spark and Mac Studio. It is available via StepFun's own platform, OpenRouter, NVIDIA NIM, and forthcoming partnerships with DeepInfra, Fireworks AI, and Modal. Where it falls short is on pure reasoning benchmarks (HLE w/ tools at 47.2% vs Claude Opus 4.7's 54.7%), terminal-heavy workflows (Terminal-Bench 2.1 at 59.5 vs GPT 5.5's 82.7), and general professional tasks (GDPVal-AA at 45.8% vs frontier models at 63%).

        분석 생성일: 2026-07-17