OpenBMB오픈소스

MiniCPM5-1B

이 모델 비교

OpenBMB 개발의 경량 고성능 모델. 1B 파라미터로 효율적인 추론을 실현합니다.

파라미터

10.8

컨텍스트

128K

라이선스

Apache 2.0

출시일

2026-05-01

일본어 처리 능력

🌐Multilingual

General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.

API 가격

이 모델의 API 가격 정보는 현재 공개되지 않았습니다

강점

    약점

      활용 사례

        심층 분석

        Parameters

        1.08B

        Dense LlamaForCausalLM architecture

        Artificial Analysis Intelligence Index

        12

        #2 among comparable open-weight models

        Context Window

        131,072 tokens

        128K native

        License

        Apache 2.0

        Fully open-source

        Price (Self-hosted)

        $0.00

        Free to run locally

        AIME 2025 (Thinking)

        40.42

        vs LFM2.5-1.2B: 31.88

        강점

        • Exceptional 1B-class SOTA in tool use, code, and reasoning tasks
        • Hybrid Think/No-Think reasoning via same checkpoint for deployment flexibility
        • Superior efficiency with up to 31x fewer output tokens than reasoning peers it surpasses

        약점

        • Limited general world knowledge and instruction following compared to larger 4B+ models
        • Performance lags on complex scientific reasoning (e.g., GPQA Diamond)
        • Slower generation speed than larger, optimized models on high-end hardware

        경쟁사 비교

        ModelArenaSWEGPQAPrice
        Qwen3-0.6B/thinkN/AN/AN/AFree (Apache 2.0)
        LFM2.5-1.2B-ThinkingN/AN/A34.85Free
        Qwen3.5-0.8B/thinkN/AN/AN/AFree

        MiniCPM5-1B is a 1.08 billion parameter dense language model from OpenBMB, released in May 2026 under the Apache 2.0 license. It is positioned as the state-of-the-art (SOTA) model within the 1B-class open-weight category, with its primary strengths in agentic tool use, code generation, and competitive mathematics. Its core innovation lies not in raw size, but in its post-training methodology: a three-stage pipeline of Supervised Fine-Tuning (SFT), Reinforcement Learning (RL), and On-Policy Distillation (OPD). This RL+OPD process trains specialized domain teachers (for math, code, QA) and distills them into a single model, yielding a +16-point average score improvement over the SFT-only baseline while drastically reducing overlong, inefficient responses.

        The model is architected for efficiency and on-device deployment. It uses the standard LlamaForCausalLM architecture, ensuring compatibility with all major inference engines (vLLM, SGLang, llama.cpp, Ollama, etc.). A key feature is its hybrid reasoning capability, toggled via a `` template and enable_thinking parameter, allowing a single checkpoint to function as both a fast conversational assistant and a deliberate reasoner. With a 128K context window and a footprint of under 1GB at Q4 quantization, it is explicitly designed for local assistants, coding agents, tool-use workflows, and resource-constrained edge scenarios where latency and token efficiency are critical. Benchmarks show it significantly outperforms models like Qwen3-0.6B and LFM2.5-1.2B in areas like AIME mathematics and τ2-Bench agentic tasks.

        분석 생성일: 2026-07-17