OpenBMB가 개발한 멀티모달 모델. 텍스트와 이미지의 통합 처리에 뛰어납니다.
파라미터
13
컨텍스트
262K
라이선스
Apache 2.0
출시일
2026-05-11
일본어 처리 능력
General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.
API 가격
이 모델의 API 가격 정보는 현재 공개되지 않았습니다
강점
약점
활용 사례
심층 분석
Parameters
1.3B
Base model: SigLIP2-400M (vision) + Qwen3.5-0.8B (LLM)
Intelligence Index
13
Surpasses Qwen3.5-0.8B (10) with 19x lower token cost
VRAM Requirement
4 GB
GGUF variant: 2 GB CPU memory
Context Window
262,144 tokens
262K tokens
Mobile Deployment
iOS, Android, HarmonyOS
Full edge adaptation code open-sourced
License
Apache 2.0
Commercial use permitted
강점
- ・Ultra-efficient architecture with >50% reduction in visual encoding FLOPs enabling real on-device inference
- ・Flexible mixed 4x/16x visual token compression allows trading accuracy for speed per task
- ・Broad mobile platform coverage (iOS, Android, HarmonyOS) with fully open-sourced edge code
약점
- ・Performance ceiling due to 1.3B parameter size limits complex multi-step reasoning tasks
- ・Real-world on-device latency and stability can vary significantly by device hardware
- ・Most benchmark results are self-reported by the OpenBMB team without independent verification
경쟁사 비교
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| Qwen3.5-0.8B | N/A | N/A | N/A | $0.00 (open weights) |
| Gemma4-E2B | N/A | N/A | N/A | $0.00 (open weights) |
| Mistral Medium 3.5 | N/A | N/A | 74.8% | $1.50/$7.50 |
MiniCPM-V 4.6 is a 1.3-billion parameter multimodal model specifically engineered by OpenBMB for ultra-efficient deployment on edge devices, particularly smartphones. Its core innovation lies in a highly optimized architecture based on LLaVA-UHD v4, which reduces visual encoding computation by over 50%, and a novel mixed 4x/16x visual token compression mechanism. This allows it to achieve superior performance-to-size ratios, scoring 13 on the Artificial Analysis Intelligence Index—outperforming the larger Qwen3.5-0.8B with drastically lower token costs.
Positioned as a 'pocket-sized MLLM', the model prioritizes practical on-device capabilities over raw cloud-scale intelligence. It supports single-image, multi-image, and video understanding, with strong results in OCR and document parsing (e.g., 84.6 on OmniDocBench). Its design philosophy is efficiency-first, enabling deployment on mainstream mobile platforms (iOS, Android, HarmonyOS) with low VRAM requirements (4GB GPU, 2GB CPU via GGUF). This makes it a compelling open-source option for developers building vision-language features directly into consumer applications.
While it cannot match the absolute performance of frontier cloud models like GPT-5 or Gemini 2.5 Flash, MiniCPM-V 4.6 carves a unique niche by bringing credible multimodal understanding to resource-constrained environments. It represents a significant step towards practical, privacy-preserving AI on personal devices, backed by a comprehensive developer ecosystem including multiple quantization formats and support for major inference frameworks like vLLM, llama.cpp, and Ollama.
출처
분석 생성일: 2026-07-17