Gemini Omni Flash is Google’s multimodal generation model, announced at I/O 2026. It takes images, audio, video, and text as input and returns video with sound — and unlike a plain text-to-video model, it edits conversationally, so each instruction builds on the last instead of starting over. These are copy-paste prompts written for that behaviour. Copy one, change the bold details, and keep the structure.
What Omni actually is
Google describes Omni as a model that can “create anything from any input — starting with video.” The parts that change how you write a prompt:
| Capability | What it means for your prompt |
|---|
| Multimodal input | Images, audio, video, and text can go in together. You must say what each input is for. |
| Conversational editing | Follow-ups build on the last result — characters, physics, and scene continuity carry over. |
| Physics reasoning | Gravity, kinetic energy, and fluid dynamics are modelled, so describe the event and let it resolve the motion. |
| World knowledge | Grounded in Gemini’s knowledge of history, science, and culture, so named specifics beat generic description. |
| Native audio | Video comes out with sound, so audio direction belongs in the prompt. |
Availability: Google AI Plus, Pro, and Ultra subscribers via the Gemini app and Google Flow, free on YouTube Shorts and the YouTube Create app, and through the API for developers. The free YouTube route is the cheapest way to try anything on this page.
The five-part structure
Every prompt here follows the same shape, and the two parts people leave out are the ones that cause the most drift:
- Goal — what you are making, how long, and how many shots. Say one continuous shot if you want one; Omni reasons about what happens next and will otherwise invent a cut.
- Input role — what each attached file is for. First frame? Style reference? Voice reference? A fidelity constraint on a product? Omni reasons across inputs together rather than stitching them in sequence, so an unlabelled input gets treated as loose inspiration.
- Scene — subject, setting, light. Named specifics outperform adjectives here more than on most models, because the world knowledge has something to attach to.
- Motion — camera and subject movement. Cinematography vocabulary works; “none” is a valid and useful answer.
- Constraints — what must not change. This is where the negatives live.
Editing is a second prompt, not a rewrite
The single biggest difference from Veo. When something is nearly right, do not re-send the whole prompt with one word changed. Send only the delta and fence the rest:
Keep everything as it is. Change only the lighting: make it late afternoon
instead of midday, warmer key from frame left, longer shadows. Do not change
the subject, the camera move, the audio, or the length.
Google’s stated caveat applies: holding complete consistency across a long chain of edits is a known limitation. Expect drift after several turns, and re-anchor with a fuller prompt when it appears.
Known limitations
Straight from Google’s model card, so you can design around them rather than discover them:
- Consistency across edits degrades over a long editing chain.
- Complex motion is still hard — busy multi-subject action is the least reliable thing to ask for.
- Accurate text rendering is unreliable. Frame lettering small, angled, or partly out of frame.
- Speech changes during editing are deliberately restricted by Google.
Coming from Veo?
Most of what you know transfers: the shot-subject-action-setting ordering, cinematography vocabulary, audio direction, and negatives all work. What is new is the input-role line and conversational editing. If you are moving between the two, the Veo prompt library covers the Veo side, and the Veo 3.1 prompting guide covers the formula these prompts are adapted from.
FAQ
What is Gemini Omni Flash?
Gemini Omni Flash is Google's multimodal generation model, announced at I/O 2026. It takes images, audio, video, and text as input and generates video with sound. Google describes it as creating "anything from any input — starting with video," with image and audio output planned later. It is a separate line from Veo, not a Veo release.
Is Gemini Omni Flash the same as Veo 4?
No. There is no Veo 4. At I/O 2026 Google's next-generation video model was announced as Gemini Omni Flash, and Veo 3.1 remains the current Veo line. Pages advertising "Veo 4 prompts" are describing a model that does not exist.
How do I access Gemini Omni Flash?
It rolled out to Google AI Plus, Pro, and Ultra subscribers through the Gemini app and Google Flow, and at no cost to users on YouTube Shorts and the YouTube Create app. Developer and enterprise access came through the API after launch. The free YouTube route is the cheapest way to try these prompts.
How is prompting Gemini Omni different from prompting Veo?
Two differences matter. First, Omni accepts image, audio, and video inputs, so a prompt has to say what each input is FOR — first frame, style reference, voice reference, or fidelity constraint. Second, Omni edits conversationally: follow-up instructions build on the previous result rather than replacing it, so a second-turn prompt should name only the change and fence everything else.
Does Gemini Omni generate audio?
Yes. It produces video with sound, so audio direction belongs in the prompt the same way it does for Veo — name the ambience, tie sound effects to visible actions, and say "no music" explicitly if you do not want a music bed.
What is Gemini Omni Flash bad at?
Google's own model card names three: holding complete consistency across a sequence of edits, scenes with complex motion, and rendering accurate text. Changing a person's speech during editing is also deliberately restricted. Frame lettering small or angled rather than fighting the text rendering.
Are these prompts verified?
No. Every prompt on this page is labelled untested. They are written to the model's documented behaviour and to the structure the format needs, but they have not been run against a real Gemini Omni output and no public evidence video is attached. Treat them as starting structures.