OpenBMBOpen Source

MiniCPM-V 4.6

Compare this model

A multimodal model developed by OpenBMB, excelling at integrated processing of text and images.

Parameters

13

Context Window

262K

License

Apache 2.0

Release Date

2026-05-11

Japanese Language Capability

🌐Multilingual

General multilingual model. Basic Japanese processing is possible, but inferior to specialized models.

API Pricing

API pricing for this model is not yet available

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        Parameters

        1.3B

        Base model: SigLIP2-400M (vision) + Qwen3.5-0.8B (LLM)

        Intelligence Index

        13

        Surpasses Qwen3.5-0.8B (10) with 19x lower token cost

        VRAM Requirement

        4 GB

        GGUF variant: 2 GB CPU memory

        Context Window

        262,144 tokens

        262K tokens

        Mobile Deployment

        iOS, Android, HarmonyOS

        Full edge adaptation code open-sourced

        License

        Apache 2.0

        Commercial use permitted

        Strengths

        • Ultra-efficient architecture with >50% reduction in visual encoding FLOPs enabling real on-device inference
        • Flexible mixed 4x/16x visual token compression allows trading accuracy for speed per task
        • Broad mobile platform coverage (iOS, Android, HarmonyOS) with fully open-sourced edge code

        Weaknesses

        • Performance ceiling due to 1.3B parameter size limits complex multi-step reasoning tasks
        • Real-world on-device latency and stability can vary significantly by device hardware
        • Most benchmark results are self-reported by the OpenBMB team without independent verification

        Competitor Comparison

        ModelArenaSWEGPQAPrice
        Qwen3.5-0.8BN/AN/AN/A$0.00 (open weights)
        Gemma4-E2BN/AN/AN/A$0.00 (open weights)
        Mistral Medium 3.5N/AN/A74.8%$1.50/$7.50

        MiniCPM-V 4.6 is a 1.3-billion parameter multimodal model specifically engineered by OpenBMB for ultra-efficient deployment on edge devices, particularly smartphones. Its core innovation lies in a highly optimized architecture based on LLaVA-UHD v4, which reduces visual encoding computation by over 50%, and a novel mixed 4x/16x visual token compression mechanism. This allows it to achieve superior performance-to-size ratios, scoring 13 on the Artificial Analysis Intelligence Index—outperforming the larger Qwen3.5-0.8B with drastically lower token costs.

        Positioned as a 'pocket-sized MLLM', the model prioritizes practical on-device capabilities over raw cloud-scale intelligence. It supports single-image, multi-image, and video understanding, with strong results in OCR and document parsing (e.g., 84.6 on OmniDocBench). Its design philosophy is efficiency-first, enabling deployment on mainstream mobile platforms (iOS, Android, HarmonyOS) with low VRAM requirements (4GB GPU, 2GB CPU via GGUF). This makes it a compelling open-source option for developers building vision-language features directly into consumer applications.

        While it cannot match the absolute performance of frontier cloud models like GPT-5 or Gemini 2.5 Flash, MiniCPM-V 4.6 carves a unique niche by bringing credible multimodal understanding to resource-constrained environments. It represents a significant step towards practical, privacy-preserving AI on personal devices, backed by a comprehensive developer ecosystem including multiple quantization formats and support for major inference frameworks like vLLM, llama.cpp, and Ollama.

        Analysis generated: 2026-07-17