Google DeepMind's Gemini Omni 1.1 Flash (Aug 27, 2026) is a multimodal video-generation and editing model that takes text, image, and video inputs and returns video plus text. Its headline Scene Extension analyzes up to 10 seconds of prior footage and chains 10-second additions to a cumulative 40 seconds, with first/last-frame control for directing motion between supplied keyframes and up to three 3-second reference clips for visual consistency. A 360p draft mode renders up to 60% faster at about one-third the cost of 720p, and output scales to 4K (upscaled). It runs on a 131,072-token input / 57,920-token output context and embeds a SynthID watermark. API pricing is $1.50/1M input, $9.00/1M text output, and about $0.10/sec of 720p video output.
Parameters
Undisclosed
Context Window
131K
License
Proprietary
Release Date
2026-08-27
API Pricing
Input Price (per 1M tokens)
$1.5
Output Price (per 1M tokens)
$9
Billing Mode: standard
Strengths
- •Scene Extension chains clips to 40 seconds with strong visual continuity
- •First/last-frame and reference-clip controls for directable video
- •360p draft mode cuts iteration cost ~3x vs 720p
- •Native multimodal input (text/image/video) with SynthID watermark
Weaknesses
- •Preview status; no system instructions, temperature, or negative-prompt controls
- •Extensions append only (no insert/prepend); audio editing unsupported
- •Full language support limited to English
Use Cases
- •Directable short-form ad and social video
- •Storyboard-to-clip prototyping in creative tools (Firefly, Figma Weave)
- •Educational and real-estate presentation videos