StepFunAIオープンソース

Step 3.7 Flash

このモデルを比較

StepFun AI開発の高速基盤モデル。低レイテンシーと高スループットを両立。

シェア:XはてブLINE

パラメータ

1980

コンテキスト長

256K

ライセンス

Apache 2.0

リリース日

2026-05-29

日本語性能

🌐多言語対応

一般的な多言語対応モデル。基本的な日本語処理は可能だが、特化モデルには劣る。

API料金

このモデルのAPI料金情報は現在未公開です

強み

    弱み

      活用例

        深度分析

        ClawEval-1.1

        67.1

        #1 overall; next best is 59.8 (Claude 4 Opus)

        SWE-Bench PRO

        56.3

        #2 overall behind Claude 4 Opus at 64.3

        SimpleVQA (Search)

        79.2

        #1; beats GPT 5.5 at 79.1

        Input Price

        $0.20/1M tokens

        $0.04/1M on cache hit; ~25× cheaper than Claude Opus

        Throughput

        400 tok/s

        cloud API; ~27 tok/s on DGX Spark locally

        Context Window

        256K tokens

        ~384 pages of text

        強み

        • Best-in-class agent reliability with ClawEval-1.1 score of 67.1, demonstrating superior multi-step tool orchestration and resistance to adversarial traps
        • Advisor Mode delivers 97% of Claude Opus 4.6 coding performance at $0.19/task vs $1.76/task — a genuine 9× cost reduction for production agentic workflows
        • Extensive local deployment flexibility: runs on Mac Studio, DGX Spark, AMD Ryzen AI Max+ 395 with BF16, FP8, NVFP4, and GGUF quantizations available on day one

        弱み

        • Significant gap on Terminal-Bench 2.1 (59.5) vs GPT 5.5 (82.7) and GDPVal-AA (45.8) vs frontier models at 63%, limiting complex terminal and general professional task use cases
        • Self-reported benchmark data; many competitor scores are from StepFun's own re-runs rather than standardized third-party evaluations, reducing comparability confidence
        • 128GB minimum hardware for local deployment excludes consumer laptops and most prosumer setups; full BF16 requires 394GB

        競合比較

        ModelArenaSWEPrice
        Claude Opus 4.6N/A64.3%$5.00/$25.00
        DeepSeek V4 FlashN/A55.6%$0.22/$0.87
        Gemini 3.5 FlashN/A55.1%$0.15/$0.60

        Step 3.7 Flash is a 198B-parameter sparse Mixture-of-Experts vision-language model released by StepFun AI on May 29, 2026. Despite its massive total parameter count, it activates only ~11B parameters per token through expert routing (8 out of 288 experts), delivering near-frontier intelligence at inference costs comparable to small dense models. The model combines a 196B language backbone with a 1.8B native vision encoder, supports a 256K context window, and achieves up to 400 tokens/second throughput on cloud APIs.

        The model's primary differentiation lies in agentic workflow reliability rather than absolute intelligence ceiling. It leads the ClawEval-1.1 benchmark at 67.1 (a 23.5-point jump over its predecessor Step 3.5 Flash) and places second on SWE-Bench PRO at 56.3. Its most innovative feature is Advisor Mode, a two-tier execution strategy where Step 3.7 Flash handles routine tool-calling and iteration while escalating to a larger frontier model only at genuine decision points, achieving approximately 97% of Claude Opus 4.6's coding performance at roughly one-ninth the per-task cost ($0.19 vs $1.76).

        The model is fully open-weight under Apache 2.0 and ships with quantized checkpoints (BF16, FP8, NVFP4, GGUF) across HuggingFace, enabling local deployment on 128GB+ unified memory devices like NVIDIA DGX Spark and Mac Studio. It is available via StepFun's own platform, OpenRouter, NVIDIA NIM, and forthcoming partnerships with DeepInfra, Fireworks AI, and Modal. Where it falls short is on pure reasoning benchmarks (HLE w/ tools at 47.2% vs Claude Opus 4.7's 54.7%), terminal-heavy workflows (Terminal-Bench 2.1 at 59.5 vs GPT 5.5's 82.7), and general professional tasks (GDPVal-AA at 45.8% vs frontier models at 63%).

        分析生成日: 2026-07-17