xAIプロプライエタリ

Grok Imagine Image Quality (2026-05-19)

このモデルを比較

xAI開発の高品質画像生成モデル。最高品質の画像出力に特化。

シェア:XはてブLINE

パラメータ

非公開

コンテキスト長

ライセンス

プロプライエタリ

リリース日

2026-05-19

日本語性能

高品質日本語

多言語対応モデルのうち、日本語処理に優れた性能を持つモデル。

API料金

このモデルのAPI料金情報は現在未公開です

強み

    弱み

      活用例

        深度分析

        Text-to-Image Arena

        Top 5

        Ranked among top models as of May 4, 2026

        Max Resolution

        2K (2048×2048)

        13 aspect ratios supported

        Text-to-Image Latency

        ~4s

        vs GPT Image 2: ~130s average

        Price (1K output)

        $0.05/image

        2K: $0.07/image

        Image-to-Video Arena

        #1 (Elo 1,336)

        Feeds xAI's #1 video pipeline

        Architecture

        Aurora MoE (Autoregressive)

        Non-diffusion; expert-specialized routing

        強み

        • Exceptionally fast inference (~4s per image) — roughly 14× faster than GPT Image 2 in head-to-head tests
        • Best-in-class image-to-video pipeline (#1 on Artificial Analysis Video Arena, Elo 1,336)
        • Strong cinematic realism with accurate skin textures, pores, subsurface scattering, and architectural composition

        弱み

        • Physics accuracy lags GPT Image 2 significantly — fluid dynamics, refraction, and material logic failures documented
        • Text rendering improved but still trails Ideogram, GPT Image 2, and FLUX; can confuse Traditional vs. Simplified Chinese
        • No multi-image reference support (single reference only); no batch generation from same prompt

        競合比較

        ModelArenaPrice
        GPT Image 2 (OpenAI)#1 (Elo 1,338)$0.009–$0.22/image
        Nano Banana Pro (Google)#3 (Elo 1,219)$0.04/image
        Grok Imagine Image Quality (xAI)Top 5$0.05–$0.07/image

        Grok Imagine Image Quality (released May 6, 2026) is xAI's flagship image generation model, built on the proprietary Aurora autoregressive Mixture-of-Experts architecture — a notable departure from the diffusion-based approach used by most competitors. It targets professional creators and enterprise teams who need photorealistic output with strong prompt adherence, accurate textures, and clean in-image typography. The model supports text-to-image generation at 1K and 2K resolution across 13 aspect ratios, plus a separate editing endpoint that enables object addition, removal, style transfer, and multi-turn iterative refinement through natural-language prompts without mask-based inpainting.

        Its most compelling differentiator is speed: independent benchmarks consistently show ~4-second latency per image, roughly 3–14× faster than competitors like GPT Image 2 and Gemini. The model also feeds directly into xAI's image-to-video pipeline, which currently ranks #1 on the Artificial Analysis Image-to-Video Arena (Elo 1,336) and Multi-Image-to-Video Arena (Elo 1,342), making it the fastest path from still image to motion content.

        However, head-to-head evaluations reveal clear tradeoffs. While Grok Imagine Image Quality excels at cinematic lighting, skin realism, and identity consistency, it struggles with complex physics (fluid dynamics, refraction), falls short of GPT Image 2 on structural editing fidelity, and currently supports only single-reference image inputs. Independent testing by Vidguru AI Lab scored it 35/45 versus GPT Image 2's 43/45 across 9 blind categories, with the largest gaps in material logic and fluid physics.

        分析生成日: 2026-07-17