StepFunAIOpen Source

Step 3.7 Flash

Compare this model

A high-speed foundational model developed by StepFun AI, balancing low latency with high throughput.

Parameters

1980

Context Window

256K

License

Apache 2.0

Release Date

2026-05-29

Japanese Language Capability

🌐Multilingual

General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.

API Pricing

API pricing for this model is not yet available

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        ClawEval-1.1

        67.1

        #1 overall; next best is 59.8 (Claude 4 Opus)

        SWE-Bench PRO

        56.3

        #2 overall behind Claude 4 Opus at 64.3

        SimpleVQA (Search)

        79.2

        #1; beats GPT 5.5 at 79.1

        Input Price

        $0.20/1M tokens

        $0.04/1M on cache hit; ~25× cheaper than Claude Opus

        Throughput

        400 tok/s

        cloud API; ~27 tok/s on DGX Spark locally

        Context Window

        256K tokens

        ~384 pages of text

        Strengths

        • Best-in-class agent reliability with ClawEval-1.1 score of 67.1, demonstrating superior multi-step tool orchestration and resistance to adversarial traps
        • Advisor Mode delivers 97% of Claude Opus 4.6 coding performance at $0.19/task vs $1.76/task — a genuine 9× cost reduction for production agentic workflows
        • Extensive local deployment flexibility: runs on Mac Studio, DGX Spark, AMD Ryzen AI Max+ 395 with BF16, FP8, NVFP4, and GGUF quantizations available on day one

        Weaknesses

        • Significant gap on Terminal-Bench 2.1 (59.5) vs GPT 5.5 (82.7) and GDPVal-AA (45.8) vs frontier models at 63%, limiting complex terminal and general professional task use cases
        • Self-reported benchmark data; many competitor scores are from StepFun's own re-runs rather than standardized third-party evaluations, reducing comparability confidence
        • 128GB minimum hardware for local deployment excludes consumer laptops and most prosumer setups; full BF16 requires 394GB

        Competitor Comparison

        ModelArenaSWEPrice
        Claude Opus 4.6N/A64.3%$5.00/$25.00
        DeepSeek V4 FlashN/A55.6%$0.22/$0.87
        Gemini 3.5 FlashN/A55.1%$0.15/$0.60

        Step 3.7 Flash is a 198B-parameter sparse Mixture-of-Experts vision-language model released by StepFun AI on May 29, 2026. Despite its massive total parameter count, it activates only ~11B parameters per token through expert routing (8 out of 288 experts), delivering near-frontier intelligence at inference costs comparable to small dense models. The model combines a 196B language backbone with a 1.8B native vision encoder, supports a 256K context window, and achieves up to 400 tokens/second throughput on cloud APIs.

        The model's primary differentiation lies in agentic workflow reliability rather than absolute intelligence ceiling. It leads the ClawEval-1.1 benchmark at 67.1 (a 23.5-point jump over its predecessor Step 3.5 Flash) and places second on SWE-Bench PRO at 56.3. Its most innovative feature is Advisor Mode, a two-tier execution strategy where Step 3.7 Flash handles routine tool-calling and iteration while escalating to a larger frontier model only at genuine decision points, achieving approximately 97% of Claude Opus 4.6's coding performance at roughly one-ninth the per-task cost ($0.19 vs $1.76).

        The model is fully open-weight under Apache 2.0 and ships with quantized checkpoints (BF16, FP8, NVFP4, GGUF) across HuggingFace, enabling local deployment on 128GB+ unified memory devices like NVIDIA DGX Spark and Mac Studio. It is available via StepFun's own platform, OpenRouter, NVIDIA NIM, and forthcoming partnerships with DeepInfra, Fireworks AI, and Modal. Where it falls short is on pure reasoning benchmarks (HLE w/ tools at 47.2% vs Claude Opus 4.7's 54.7%), terminal-heavy workflows (Terminal-Bench 2.1 at 59.5 vs GPT 5.5's 82.7), and general professional tasks (GDPVal-AA at 45.8% vs frontier models at 63%).

        Analysis generated: 2026-07-17