xAI開発の高品質画像生成モデル。最高品質の画像出力に特化。
パラメータ
非公開
コンテキスト長
ライセンス
プロプライエタリ
リリース日
2026-05-19
日本語性能
多言語対応モデルのうち、日本語処理に優れた性能を持つモデル。
API料金
このモデルのAPI料金情報は現在未公開です
強み
弱み
活用例
深度分析
Text-to-Image Arena
Top 5
Ranked among top models as of May 4, 2026
Max Resolution
2K (2048×2048)
13 aspect ratios supported
Text-to-Image Latency
~4s
vs GPT Image 2: ~130s average
Price (1K output)
$0.05/image
2K: $0.07/image
Image-to-Video Arena
#1 (Elo 1,336)
Feeds xAI's #1 video pipeline
Architecture
Aurora MoE (Autoregressive)
Non-diffusion; expert-specialized routing
強み
- ・Exceptionally fast inference (~4s per image) — roughly 14× faster than GPT Image 2 in head-to-head tests
- ・Best-in-class image-to-video pipeline (#1 on Artificial Analysis Video Arena, Elo 1,336)
- ・Strong cinematic realism with accurate skin textures, pores, subsurface scattering, and architectural composition
弱み
- ・Physics accuracy lags GPT Image 2 significantly — fluid dynamics, refraction, and material logic failures documented
- ・Text rendering improved but still trails Ideogram, GPT Image 2, and FLUX; can confuse Traditional vs. Simplified Chinese
- ・No multi-image reference support (single reference only); no batch generation from same prompt
競合比較
| Model | Arena | Price |
|---|---|---|
| GPT Image 2 (OpenAI) | #1 (Elo 1,338) | $0.009–$0.22/image |
| Nano Banana Pro (Google) | #3 (Elo 1,219) | $0.04/image |
| Grok Imagine Image Quality (xAI) | Top 5 | $0.05–$0.07/image |
Grok Imagine Image Quality (released May 6, 2026) is xAI's flagship image generation model, built on the proprietary Aurora autoregressive Mixture-of-Experts architecture — a notable departure from the diffusion-based approach used by most competitors. It targets professional creators and enterprise teams who need photorealistic output with strong prompt adherence, accurate textures, and clean in-image typography. The model supports text-to-image generation at 1K and 2K resolution across 13 aspect ratios, plus a separate editing endpoint that enables object addition, removal, style transfer, and multi-turn iterative refinement through natural-language prompts without mask-based inpainting.
Its most compelling differentiator is speed: independent benchmarks consistently show ~4-second latency per image, roughly 3–14× faster than competitors like GPT Image 2 and Gemini. The model also feeds directly into xAI's image-to-video pipeline, which currently ranks #1 on the Artificial Analysis Image-to-Video Arena (Elo 1,336) and Multi-Image-to-Video Arena (Elo 1,342), making it the fastest path from still image to motion content.
However, head-to-head evaluations reveal clear tradeoffs. While Grok Imagine Image Quality excels at cinematic lighting, skin realism, and identity consistency, it struggles with complex physics (fluid dynamics, refraction), falls short of GPT Image 2 on structural editing fidelity, and currently supports only single-reference image inputs. Independent testing by Vidguru AI Lab scored it 35/45 versus GPT Image 2's 43/45 across 9 blind categories, with the largest gaps in material logic and fluid physics.
出典
- Grok Imagine Quality Mode API | xAI Official Announcement
- Grok Imagine Image Quality API Documentation & Pricing | Atlas Cloud
- Grok Imagine Image Quality Model Page | AI Stats
- Grok Imagine Quality vs GPT Image 2: Blind Test Comparison | Vidguru
- Grok Imagine Quality Mode API Release Analysis | CometAPI
- Grok Imagine Image vs GPT Image 2: 6 Category Benchmark | Atlas Cloud
- Grok Imagine Image Benchmark: 2K Start Frames | GenFlick
- Grok Imagine Image vs GPT Image 1.5: 5 Prompt Tests | Wiro.ai
- Grok Imagine Review: Image and Video Inside Grok | Vuela.ai
- Grok Imagine Image Quality API Pricing | OpenRouter
分析生成日: 2026-07-17