How to Prompt Gemini Omni Flash 1.1 Video Model
Learn how to prompt Gemini Omni Flash 1.1 for text-to-video, image-to-video, first and last frames, reference images, video editing, video extension, audio, timing, and text.
Omni generates sound with every video, builds multi-shot sequences by default, and lets you edit through conversation. The mistake is treating it like a one-shot generator.
If you write one prompt, generate, and start over when the result is wrong, you are doing it the hard way. Omni remembers what it made. You can tell it to change one thing and keep everything else.
Audio is the other thing people miss. Every video comes with sound. If you do not describe what you want to hear, Omni decides for you. Sometimes that works. Sometimes you get dialogue you did not ask for.
If you want one continuous shot, you have to ask for it. Otherwise Omni cuts the video into several angles on its own.
What Google Omni Is
Omni Flash launched as a public preview on June 30, 2026. Omni Flash 1.1 became generally available on August 27, 2026. The preview endpoint (gemini-omni-flash-preview) turns off on September 30, 2026.
Both versions share the same core. Text, image, audio, and video go in. 3 to 10 seconds per generation. 16:9 or 9:16. 24 frames per second. World knowledge that fills in physics, history, and culture without you describing them.
What 1.1 adds
- First and Last Frames. Two images define the start and end of the video. The model generates the footage between them. This did not exist in the preview.
- Video Extension. Append up to 10 seconds at the end of a clip, up to 40 seconds total. The model reads the last 10 seconds of context to keep the continuation consistent.
- Resolution selector. 360p, 720p, 1080p, and 4K. 720p is the default. 1080p and 4K are upscaled from the 720p render, not native generations. 360p is a draft tier that generates faster and costs less.
- Free-form durations. Any integer from 3 to 10 seconds. The preview only allowed 3, 5, or 10.
- Reference Images. Bind images to roles inside the prompt with
<IMAGE_REF_0>,<IMAGE_REF_1>, and so on. The preview accepted a bare list of reference images with no way to tell the model which one controls what.
What changed
- Reference images. The preview accepted up to 14. Version 1.1 accepts up to 10, but with the tag system you can assign each one a specific role.
- Reference video. Supported in both versions. Up to 3 clips, 3 seconds each. Audio in reference video is ignored.
First and last frame control, video extension, reference image tags, and free-form durations are 1.1 features. Everything else in this guide applies to both versions.
The Core Prompt Structure
Five elements give you control over Omni's output.
The five parts
- Shot framing and motion. Wide, medium, or close-up. How the camera moves. Should it glide, rush, or stay still? For a full breakdown of shot types, focal lengths, and camera language, see the Camera Shots Prompt Guide.
- Style. Realistic, cinematic, grounded, majestic, anime, editorial. Tell Omni the effect you want.
- Lighting. Where the light comes from. The sun, a streetlamp, off-screen. Crisp, warm, or ethereal. For a deep dive on lighting styles and how to describe them, see the Lighting Prompt Guide.
- Location. The setting. "An alien landscape with clear, azure water."
- Action. What is happening. Who the characters are. How they move and interact.
Where audio goes
Audio goes at the end of the prompt. Use a label like Sound design: followed by specific sources.
Sound design: the faint tick of the escapement, a small click of metal on metal, quiet workshop room tone. No music.Name real sound sources. "Cinematic audio" or "ambient sound" gives Omni nothing specific to work with. "Wind buffeting the hull, one distant foghorn, low sub-bass hum" gives the model real sources to generate.
To remove audio you do not want:
No dialogue.
No extra sound effects.
No music.A simple text-to-video prompt
Peaceful yellow flower field at sunrise. Start with a close-up of one flower covered in morning dew, gently moving in the breeze. Without warning, a large agricultural harvesting machine suddenly comes through the flowers and cuts directly through the flower in the foreground. The camera quickly pulls back and rises, revealing the machine harvesting a long row of flowers across a huge commercial field. Show realistic cutting machinery, petals and stems moving naturally, untouched flowers ahead of the machine, and a clearly harvested strip behind it. Photorealistic, cinematic, natural morning light, realistic plant physics, realistic machinery, one continuous shot, no scene changes, no text or logos.- Shot framing and motion: "Start with a close-up of one flower... The camera quickly pulls back and rises... one continuous shot, no scene changes"
- Action: "A large agricultural harvesting machine suddenly comes through the flowers and cuts directly through the flower in the foreground... harvesting a long row of flowers across a huge commercial field"
- Lighting: "Natural morning light"
- Location: "Peaceful yellow flower field at sunrise... a huge commercial field"
- Style: "Photorealistic, cinematic, realistic plant physics, realistic machinery"
- Preservation: "No text or logos"
Style is included here as "photorealistic, cinematic." If you want a different look, swap it: "anime style," "claymation," "watercolor." If you omit style entirely, Omni defaults to a photorealistic look.
Controlling the Defaults
One continuous shot
Omni builds multi-shot narratives by default. To force a single unbroken shot, use any of these phrases:
In a single unbroken scene
In a single continuous shot
One continuous shot
No scene cutsOner also works. It is a film term for a single continuous shot with no cuts.
Use one continuous shot when the action is simple and the camera movement carries the scene. Use multiple shots when you want Omni to build a narrative with different angles.
Camera vocabulary
Omni responds to specific camera terms.
Static angles:
staticorlocked offorfixed: camera does not moveeye-level: camera at the subject's eye heightoverhead: camera directly above the subject
Movement:
push in: camera moves slowly toward the subjectpunch in: fast push toward the subjectdolly zoom: camera moves forward while zooming out, or vice versaorbit: camera circles the subjecttrack: camera moves alongside the subjecthandheld: camera held by hand, slight shake and driftdrone: aerial camera movement
Camera types:
natural smartphone zoom: the look of a phone camera zoomingfilm camera: the texture and feel of analog filmwebcam style: the look of a webcam feed
Subject motion and camera motion are separate. "A woman walking" describes what the subject does. "Tracking shot alongside a woman walking" describes what the camera does. Specify both when both matter. For the full vocabulary of shot types, focal lengths, and framing, see the Camera Shots Prompt Guide.
World knowledge: when to explain, when to stop
You do not need to describe everything. If Omni already understands the subject, the setting, or the physics, name it and move on. "A samurai in feudal Japan" is enough. The model knows the clothing, architecture, and setting.
Be specific when the result must match an exact visual. If you need a particular color, a particular angle, or a particular material, say so. If you just need "a cozy mountain kitchen," Omni can fill in the rest.
Timing and Text
Timing events
You can stage events at specific points in the video.
Natural language:
After 3 seconds, a woman enters the scene.
At 5s a door slams in the background audio.
Every 2s cut to a new shot.Timecode syntax:
[0-3s] A person is walking
[3-6s] They stop and turn around
[6-10s] They start runningUse natural language for one or two timed events in a scene. Use timecodes when you need to script the entire clip from start to finish.
Text in videos
Short text works. Long text degrades. Put the exact text in quotation marks and limit it to a few words.
Create an explosive, motion-graphics-driven character spot in 16:9, exactly 10 seconds total, 24fps. Hero: animated girl in blue pajamas. Preserve exact design: soft curly dark hair, yellow hair clip, warm brown skin, expressive brown eyes, blue polka-dot t-shirt, white over-ear headphones around neck, blue patterned pants, stylized 3D proportions. Never redesign or restyle. 70% bold graphics, 30% character performance. Use kinetic typography, geometric shapes, color-bursts, halftone, motion trails, confetti bursts. Graphics react to her mood. TYPOGRAPHY RULE: ONLY words allowed are CURIOUS, PLAYFUL, BRIGHT, ME. No other text, logos, or captions. Palette: sky blue, cream/white, sunshine yellow, pure white. Black for lines/type. Arc: CURIOUS > SILLY > CONFIDENT > JOYFUL. Each beat clear on face/posture. CUTS: 0.00-0.75s CU eyes tilt, CURIOUS flashes. 0.75-1.50s pull back, shapes orbit. 1.50-2.30s silly face, PLAYFUL wobbles. 2.30-3.10s triptych silly faces. 3.10-3.90s confident, BRIGHT holds. 3.90-4.70s BRIGHT huge, she steps through letters. 4.70-5.50s joyful laugh, confetti bursts. 5.50-6.30s rapid montage all 4 expressions. 6.30-7.10s spin with motion streaks. 7.10-7.90s freeze smile, ME appears. 7.90-8.70s hero rings, BRIGHT+ME hover. 8.70-9.40s final lockup with all 4 words. 9.40-10.00s pulse to black. EDITING: hard cuts, snap zooms, shutter flashes timed to expressions. CAMERA: warm, friendly, eye-level. LIGHTING: storybook-bright. BGM: upbeat electronic-acoustic, plinks, bass, giggle sting, chime, stinger at 10s. NEGATIVE: no extra text, no redesign, no distortions.This prompt controls text through three mechanisms:
- The exact words are named in uppercase:
CURIOUS,PLAYFUL,BRIGHT,ME. The model treats these as the only text allowed on screen. - A typography rule locks it down:
TYPOGRAPHY RULE: ONLY words allowed are CURIOUS, PLAYFUL, BRIGHT, ME. No other text, logos, or captions. - Timecodes schedule when each word appears:
0.00-0.75s CU eyes tilt, CURIOUS flashestells the model exactly when to show each word and what the character is doing at that moment.
The NEGATIVE: no extra text, no redesign, no distortions line reinforces the constraint from the other side.
Good uses for text in videos:
- Signs and labels: a storefront sign reading "OPEN"
- Title cards: "Chapter 1" at the start
- Captions: a single line of dialogue at the bottom
- Product text: a brand name on a package
- Kinetic typography: bold words flashing in sync with action
Keep text to a few words. Full sentences and paragraphs will not render cleanly. If you need more text, composite it in post.
Image-to-Video
When you start from an image, the image already carries appearance, composition, lighting, color, and style. The prompt does not need to rebuild any of that from text.
The prompt adds what the image cannot provide: motion, camera behavior, sound, and anything that must stay consistent during the motion.
Create a 10-second premium fashion eyewear advertisement. The uploaded image must remain the primary reference for: the model's appearance, sunglasses shape and color, WMP EYEWEAR branding, product proportions, luxury advertising aesthetic.What to put in the prompt
- Motion: Describe what moves and how. "The cat's tail twitches slowly" is useful. "Make it move" is not.
- Camera: Use the camera vocabulary from the previous section. "Slow push in" or "static, locked off."
- Sound: Use the audio format from the core prompt section.
Sound design:followed by specific sources. - Preservation: If something in the image must not change, say so. "Keep the lighting exactly as in the image."
What not to put in the prompt
- The subject's appearance. The image already defines it.
- The composition. The image already defines it.
- The lighting. The image already defines it, unless you want it to change during the video.
Image quality
Use a high-resolution image. Low-resolution inputs produce low-resolution motion.
First and Last Frames
You provide two images. The first becomes the opening frame. The second becomes the closing frame. The model generates the footage that connects them. You are giving it two fixed points and asking it to solve for the middle.
This is different from image-to-video, where one image defines the start and the model decides where the scene ends. With first and last frame control, you control both endpoints. The model's job is narrower: fill the gap between two known states.
What the prompt controls
The prompt does not re-describe what the images already show. The first image already defines the starting appearance, composition, and lighting. The last image already defines the ending appearance, composition, and lighting. The prompt controls three things:
- The progression between the two states. What happens step by step. What changes first, what changes second, what arrives last.
- Camera behavior during the transition. Should the camera stay fixed, move, orbit, or push in? If the camera should not move, say so explicitly.
- Shot structure. One continuous shot or multiple cuts. If you want one unbroken take, say "one continuous shot" or "no scene cuts."
Example
First Frame:
Last Frame:
Output:
One continuous shot from a fixed tripod position, no camera drift. The scene begins at the first frame and progresses naturally to the last frame. Workers build the pool area step by step: the decking goes down, stone edges are laid, water fills in, and furniture appears in believable stages. Materials arrive and are placed progressively. The scene finishes exactly on the completed poolside shown in the last frame.The first image shows an unfinished pool area. The last image shows the completed poolside. The prompt does not describe what either image looks like. It describes the progression: decking, stone edges, water, furniture. It locks the camera to a fixed position. It requires one continuous shot with no cuts. And it tells the model to land exactly on the last frame.
When to use it
- Camera moves that must land on a specific composition: orbits, zoom transitions, reveals
- Loops where the last frame matches the first
- Product reveals where you control both the starting and ending state
- Match cuts between two scenes
- Construction or transformation sequences where you want to control both the before and after
This moves storyboard control from the expensive video step to the cheap image step. You art-direct two stills in an image model, then pay for video once the endpoints are approved.
Tags
Use <FIRST_FRAME> and <LAST_FRAME> tags when you need to be explicit about which image is which.
First and last frame control cannot be combined with reference images in the same call. You either use frame control or reference images, not both.
Reference Images
How reference images differ from first frames
A first frame becomes the opening of the video. A reference image guides the model on what a character, object, style, or environment should look like, without appearing as a literal frame.
Binding references with tags
Place <IMAGE_REF_0>, <IMAGE_REF_1>, and so on inside the prompt where you want each reference to apply. The number in the tag matches the order of the images you upload. <IMAGE_REF_0> is the first image, <IMAGE_REF_1> is the second.
Two things make a multi-reference prompt work:
- An opening line that binds each tag to its role. Tell the model what each image is before describing the action. This prevents the model from guessing which image controls what.
- Tags placed where each reference matters in the scene. Put
<IMAGE_REF_0>where the subject appears. Put<IMAGE_REF_1>where the object or environment appears.
A 9:16 vertical product video ad, 10 seconds, 24fps. <IMAGE_REF_0> is a cool Asian woman in a navy blue No.54 football jersey and NY cap. <IMAGE_REF_1> is a dark green vintage car.
[0-3s] Fixed wide shot, 35mm film texture with fine film grain. The woman walks into the frame in front of a stone building where the car is parked, leaning effortlessly on the hood. Ultra-thin elegant typography "NAVY VIBES" at the top. Low-contrast cinematic color grading with lifted blueish shadows. Sound design: punchy hip-hop beats with subtle film projector static.
[4-7s] Close-up with organic handheld movement, smooth push-in to her face. She wears star-embellished Y2K sunglasses and touches a silver cross necklace. Soft light reflects off the jersey's printed numbers and silver waist chain. Ultra-thin artistic font "STREET ICON" on the left. Sound design: subtle metallic clinking of chains synchronized with the beat.
[8-10s] Medium-wide pull-back shot. The woman carries a denim shoulder bag and walks away toward the street corner. Smooth parallax retreat. The navy blue outfit contrasts with the green car finish. Minimalist thin font "LOOKBOOK" at the bottom center. Sound design: the beat intensifies, then cuts to a low-frequency ambient echo.This prompt uses two references for two different jobs. The opening line binds <IMAGE_REF_0> to the subject and <IMAGE_REF_1> to the car. The timecode blocks then place each tag where it matters: the woman appears in every shot, the car appears in the wide shots and the pull-back. The text overlays ("NAVY VIBES", "STREET ICON", "LOOKBOOK") are short, quoted, and placed at specific screen positions. Sound design is specified per segment instead of one generic audio instruction.
What references are good at
- Subject identity. Hairstyle, facial features, clothing details, and scars carry from a studio still into a completely different scene.
- Style consistency. Pass a reference image in a particular art style and the output follows it.
- Product consistency. A product shot used as a reference keeps the product looking the same across different scenes.
What references struggle with
Action fidelity. The model renders the subject well but follows motion instructions loosely. If you ask for someone walking with hands in pockets while a handheld camera tracks beside them, you may get someone standing still. Write prompts that lean on atmosphere and keep the motion simple.
Editing Through Conversation
You generate a video, then send a follow-up prompt that edits it. You do not re-upload anything. The model holds the previous state and applies only the change you describe.
Simple prompts work better
One change per turn. The edit prompt should be shorter than the original generation prompt. You are not writing a new scene. You are pointing at one thing and telling the model what to do with it.
Before:
After:
Change the butterfly to a bee.That is the entire edit prompt. No description of the scene, no camera instructions, no lighting, no audio. The model already has all of that from the original generation. It changes the butterfly to a bee and keeps everything else.
Add "Keep everything else the same" when the rest of the video must not move:
Make the phone invisible. Keep everything else the same.Long edit prompts cause unintended changes. The more you describe, the more the model rewrites. If you want to change one thing, say one thing.
What you can edit
- Subjects: "Change the butterfly to a bee"
- Camera: "Change the camera angle to be over the violinist's shoulder"
- Action: "The lights of the apartments start turning on in sync with the music"
- Style: "Make this video anime"
- Background: "Change the ships to be made from white origami paper"
Editing your own videos
You can upload a video and edit it. Upload up to 10 seconds.
Upload and editing are unavailable in the European Economic Area, Switzerland, and the UK. Model-generated videos remain editable everywhere.
Video Extension
Extension appends footage to the end of a clip only. You cannot prepend footage or splice into the middle. Each extension adds up to 10 seconds. The model reads the last 10 seconds of the original clip as context to keep the continuation consistent.
How to prompt an extension
The extension prompt describes what happens next. It does not re-describe the scene. The model already has the original clip. You are telling it where to go from there.
Original:
Extended:
The original video was generated from this prompt:
A cinematic tracking shot follows a Black woman walking with confidence down a path with plants on either side. With every stride, the vines from the plants continue growing, following her momentum, eventually growing towards her and onto her body. The camera cuts to a close-up and moves upwards as vines start growing from her feet, forming a pair of heels, then continues upwards wrapping around her body, forming a vine-like dress with iridescent leaves. The camera moves to her face as she looks down with a subtle smirk. She stops walking, analyzes her dress, and says "Hmmm, I like it. Nice one nature!" Sound design: natural rustling of vines growing, soft footsteps on the path, her voice. No music.The extension prompt:
Extend this video. She pauses and says "Although I think you forgot one thing." The shot cuts to her hands as roots start forming a beautiful handbag in her hands, representative of the vines and iridescent leaves. The handbag forms fully. She looks at it and says "Perfect!" She showcases the handbag. Sound design: the rustling continues as the handbag forms, her voice. No music.The extension prompt does three things:
- Opens with "Extend this video." This tells Omni to append to the previous clip rather than generate a new one.
- Describes only the new action. The woman, the dress, the setting, and the vine aesthetic are already established in the original clip. The extension prompt does not re-describe any of it. It introduces one new element: the handbag forming from roots.
- Carries the audio forward. The rustling continues from the original clip, and her voice carries over.
No musicis repeated so Omni does not introduce a soundtrack in the extension.
Audio continuity
Audio carries from the original clip unless you say otherwise. If the original had ticking, rain, or music, it keeps playing. To change the audio in the extension, describe what should happen:
The ticking continues, softer now. A new sound: the creak of a door opening.The ticking fades out. Only the distant sound of evening street traffic remains.Extension with reference media
You can add reference images to introduce new characters or objects in the extension:
Extend this video. A customer <IMAGE_REF_0> enters the frame from the right, places a broken pocket watch on the bench, and waits. The watchmaker looks up. Sound design: the door creak, footsteps on wood, the click of the watch placed on the bench. No music.You cannot add dialogue when extending an uploaded video.
Common Prompt Problems
| Problem | First thing to change |
|---|---|
| Video has unwanted cuts | Add "one continuous shot" or "no scene cuts" |
| Audio is wrong or generic | Name specific sound sources, or add "No dialogue" |
| Image-to-video motion is vague | Describe the exact motion, not "make it move" |
| References get mixed up | Use <IMAGE_REF_0> tags to bind each image to a role |
| Edit changes too much | Simplify to one change, add "Keep everything else the same" |
| Video ends too soon | Use extension: "Extend this video" |
| Text in video is garbled | Keep text to a few words, put it in quotes |
| Multi-shot sequence is chaotic | Use timecode syntax [0-3s] |
| Subject changes across shots | Add a reference image for the subject |
| Scene does not feel realistic | Let Omni's world knowledge fill in details instead of over-describing |
| First and last frame not working | This is a 1.1 feature. It does not work on the preview version |
| Extension not working | This is a 1.1 feature. It does not work on the preview version |
Before You Generate
Which resolution should I use for Google Omni Flash 1.1
720p is where Omni actually renders. 1080p and 4K are upscaled from that 720p render. 360p is a separate draft tier that generates faster and costs less.
The practical workflow: draft at 360p while you iterate on the prompt. Once the prompt, references, and composition are locked, render the final at 720p. Upscale to 1080p or 4K only for delivery. Do not pay for 4K during iteration.
What aspect ratios does Google Omni Flash support
16:9 for landscape: product demos, cinematic shots, anything viewed on a desktop or TV. 9:16 for vertical: social media, Reels, Shorts, Stories. The prompt's composition should match the ratio you select. A wide establishing shot described in a 9:16 prompt will be cropped.
Does Google Omni Flash support non-English prompts
English is fully supported. Other languages are untested. Prompts written in English will produce the most reliable results.
Does Google Omni Flash watermark its videos
Every Omni output carries a SynthID watermark. It is invisible to viewers and detectable programmatically. Plan for this if your workflow makes provenance claims.
What content does Google Omni Flash refuse to generate
Omni applies content safety filters to both input prompts and generated video. These filters vary by region. Uploading and editing images containing certain recognizable people is not supported. Uploading and editing images containing minors is not supported in the European Economic Area, Switzerland, and the United Kingdom. Prompts that violate usage policies are blocked.
How to use reference video with Google Omni Flash
Reference video is supported in both versions. Upload up to 3 clips, up to 3 seconds each. Audio in reference video is ignored. Video references work best with likenesses. Referencing or reasoning across multiple videos is not supported and may result in degraded performance.
Use the <VIDEO_REF_0> tag to bind a reference video to a role in the prompt, the same way you use <IMAGE_REF_0> for reference images.
Final Takeaway
The biggest shift with Omni is workflow, not prompt length. Generate early, edit often, and let the model hold the state between turns.
Last updated on
How to Prompt Nano Banana 2
A complete guide to writing prompts for Nano Banana 2, covering the six-part brief, camera, lighting, style, materials, text, references, image editing, and real-time images.
PhotoGPT Advanced Settings for Image Generation
In this guide, we'll go through the advanced options available in your dashboard for higher control over your generated images.