Google has launched Gemini Omni 1.1 Flash, a production-ready video generation model that gives developers and creators frame-level control over AI video. Announced on August 27, 2026, it extends scenes up to 40 seconds, renders at up to 4K, and undercuts most rival video APIs on price.

What You Can Do With It Today

Open Google AI Studio, select the Omni 1.1 Flash model, and generate a clip from a text prompt or a starting image. To carry a character or motion style across shots, upload up to three seconds of reference footage. Set a first and last frame to lock a camera move between two keyframes, then extend the scene in 10-second increments up to 40 seconds. Developers can wire the same controls into an app through the Gemini API, drafting at 360p for speed before upscaling the final cut to 1080p or 4K.

Why It Matters for Creators

Longer, controllable shots are the missing piece in most text-to-video tools. Omni 1.1 Flash analyzes 10 seconds of prior footage, up from one second, when extending a scene, which keeps motion and characters consistent across a 40-second sequence. The draft-then-upscale flow also cuts cost: 360p drafting runs up to 60 percent faster at roughly a third of the 720p price, so creators can iterate cheaply and pay for high resolution only on the final render.

Key Details

Resolutions: 360p, 720p, 1080p, and 4K, with per-second pricing of $0.03, $0.10, $0.15, and $0.30 respectively (see the Gemini API pricing).

Scene extension: Up to 40 seconds, in 10-second increments, using 10 seconds of prior context for consistency.

Controls: First and last frame selection, plus up to 3 seconds of video reference for style and character continuity.

Availability: Google AI Studio and the Gemini API now, plus Google Flow for AI Plus, Pro, and Ultra subscribers and a scene-extension feature in the Gemini app.

What to Do Next

If you already build AI video pipelines, benchmark Omni 1.1 Flash against your current tool on a 30-second shot, the length where consistency usually breaks down. Our guides on native 30-second video in ComfyUI and turning a blog post into a video show where a controllable long-form model like this fits into a working workflow.