If you already write good Veo prompts, most of that knowledge moves to Gemini Omni Flash unchanged. Shot composition, subject specificity, cinematography vocabulary, audio direction, negatives — all of it still works. Two things are genuinely different, and both are about what Omni can do that Veo cannot: it takes files as input, and it edits conversationally. This is what changes and what does not.
First: there is no Veo 4
Worth clearing up, because a lot of pages competing for these searches get it wrong. At Google I/O 2026 the next-generation video model was announced as Gemini Omni Flash — not Veo 4. Veo 3.1, along with its Fast and Lite variants, remains the current Veo line. If you have been waiting for Veo 4 to arrive before learning a new prompt format, the thing that actually arrived is Omni.
What the two models are
| Veo 3.1 | Gemini Omni Flash | |
|---|---|---|
| Input | Text (and image, for image-to-video) | Text, images, audio, video — together |
| Output | Video, 720p or 1080p, native synced audio | Video with audio; image and audio output planned |
| Editing | Re-prompt from scratch | Conversational — each instruction builds on the last |
| Strengths | Stable camera work, clean physics, mature prompt conventions | Multimodal input, physics reasoning, world knowledge, editing |
| Free route | Veo 3.1 Lite on any Google account | YouTube Shorts and YouTube Create app |
| Paid route | Veo 3.1 / 3.1 Fast | Google AI Plus, Pro, Ultra via Gemini app and Google Flow |
What transfers unchanged
Structure. Veo’s five-part formula — shot composition, subject, action, setting, mood — maps cleanly onto Omni’s scene and motion sections. Front-loading the important elements still matters on both.
Specificity. “A matte-black ceramic mug” beats “a nice mug” on both models. On Omni it matters slightly more, because the model is grounded in Gemini’s world knowledge and named specifics give that knowledge something to attach to. A prompt naming a real period, place, or process gets more out of Omni than a generic description does.
Cinematography vocabulary. Dolly, orbit, crane, locked-off, shallow depth of field — both models were trained on footage discussed in these terms, and both respond to them more reliably than to plain description.
Audio direction. Both generate sound with the picture. Name the ambience, tie sound effects to visible actions, and exclude music explicitly if you do not want a music bed. This carries over one-for-one.
Negatives. Phrase them as descriptions of absence rather than commands on both. “No on-screen captions, no background music, no zoom” works the same way in each.
What you have to write differently
1. Every input needs a role
This is the big one. Omni reasons across all inputs together rather than processing them in turn, which is powerful and also means an unlabelled attachment gets treated as vague inspiration. The same image can be four completely different instructions:
Input role: use the attached image as the exact first frame — keep the
subject, framing, and colour grade unchanged.
Input role: use the attached image as a style reference only — match its
palette and grain, not its content.
Input role: the attached image is the product and its exact colour and
finish; do not restyle it.
Input role: use the attached audio only as a voice reference for tone and
pacing — do not reuse the words.
Veo’s image-to-video has one implicit mode: the image is the starting frame. Omni has many, and it will not guess correctly on your behalf. If you take one habit from this guide, take this one.
2. Edits are deltas, not rewrites
On Veo, fixing one thing means re-sending the whole prompt with that thing changed, and accepting that everything else will re-roll too. On Omni, the scene persists — Google’s phrasing is that “every instruction builds on the last,” with characters, physics, and continuity carried forward.
So the follow-up prompt should name only the change and explicitly fence the rest:
Keep everything as it is. Change only the lighting: make it late afternoon
instead of midday, warmer key from frame left, longer shadows. Do not change
the subject, the camera move, the audio, or the length.
Sending a full rewritten prompt as turn two throws away the thing that makes Omni worth using.
The caveat is Google’s own: holding complete consistency across a long chain of edits is a documented limitation. Expect drift after several turns and re-anchor with a full prompt when you see it.
3. Shot count has to be stated
Omni reasons about what should happen next, which is useful for storytelling and inconvenient when you wanted one clean shot. It will add its own cuts if you leave the shot count open. one continuous shot, no cuts belongs in most Omni prompts and is rarely needed on Veo.
Where each one wins
Reach for Veo 3.1 when you want a single well-controlled shot from a text description, when camera stability matters, or when you are iterating cheaply on Fast before a final render. The prompt conventions are more mature and there is more public evidence of what works.
Reach for Gemini Omni when you are starting from material you already have — a photo, a clip, a voice — when you want to refine a result through several passes rather than re-roll it, or when the shot depends on physical behaviour like fluid, weight, or momentum.
Neither is good at rendering accurate text or handling busy multi-subject motion. Frame lettering small or angled on both.
Practical next steps
Copy-paste prompts for each are in the Gemini Omni prompt library and the Veo prompt library. For the Veo formula in full, see the Veo 3.1 prompting guide; for the structural parts both models share, the prompt formula and structure guide and the camera movement guide apply to Omni as well.