Gemini Omni Prompts

Gemini Omni Flash is Google’s multimodal generation model, announced at I/O 2026. It takes images, audio, video, and text as input and returns video with sound — and unlike a plain text-to-video model, it edits conversationally, so each instruction builds on the last instead of starting over. These are copy-paste prompts written for that behaviour. Copy one, change the bold details, and keep the structure.

What Omni actually is

Google describes Omni as a model that can “create anything from any input — starting with video.” The parts that change how you write a prompt:

CapabilityWhat it means for your prompt
Multimodal inputImages, audio, video, and text can go in together. You must say what each input is for.
Conversational editingFollow-ups build on the last result — characters, physics, and scene continuity carry over.
Physics reasoningGravity, kinetic energy, and fluid dynamics are modelled, so describe the event and let it resolve the motion.
World knowledgeGrounded in Gemini’s knowledge of history, science, and culture, so named specifics beat generic description.
Native audioVideo comes out with sound, so audio direction belongs in the prompt.

Availability: Google AI Plus, Pro, and Ultra subscribers via the Gemini app and Google Flow, free on YouTube Shorts and the YouTube Create app, and through the API for developers. The free YouTube route is the cheapest way to try anything on this page.

The five-part structure

Every prompt here follows the same shape, and the two parts people leave out are the ones that cause the most drift:

  1. Goal — what you are making, how long, and how many shots. Say one continuous shot if you want one; Omni reasons about what happens next and will otherwise invent a cut.
  2. Input role — what each attached file is for. First frame? Style reference? Voice reference? A fidelity constraint on a product? Omni reasons across inputs together rather than stitching them in sequence, so an unlabelled input gets treated as loose inspiration.
  3. Scene — subject, setting, light. Named specifics outperform adjectives here more than on most models, because the world knowledge has something to attach to.
  4. Motion — camera and subject movement. Cinematography vocabulary works; “none” is a valid and useful answer.
  5. Constraints — what must not change. This is where the negatives live.

Editing is a second prompt, not a rewrite

The single biggest difference from Veo. When something is nearly right, do not re-send the whole prompt with one word changed. Send only the delta and fence the rest:

Keep everything as it is. Change only the lighting: make it late afternoon
instead of midday, warmer key from frame left, longer shadows. Do not change
the subject, the camera move, the audio, or the length.

Google’s stated caveat applies: holding complete consistency across a long chain of edits is a known limitation. Expect drift after several turns, and re-anchor with a fuller prompt when it appears.

Known limitations

Straight from Google’s model card, so you can design around them rather than discover them:

Coming from Veo?

Most of what you know transfers: the shot-subject-action-setting ordering, cinematography vocabulary, audio direction, and negatives all work. What is new is the input-role line and conversational editing. If you are moving between the two, the Veo prompt library covers the Veo side, and the Veo 3.1 prompting guide covers the formula these prompts are adapted from.

Prompt deck

Copy a format, check the evidence, then customize it.

14 prompts 0 evidenced 0 community 0 owner-tested

Gemini Omni Flash / All use cases

Text-to-video — the baseline five-part prompt

Prompt
Goal: a single continuous 8-second establishing shot, no cuts. Scene: a **weathered fishing boat** moored at a **stone harbour wall at dawn**, mist on the water, ropes slack. Motion: slow dolly-in from a wide shot to a medium on the bow, camera at deck height. Audio: water lapping against the hull, distant gulls, a rope creaking under tension. No music, no dialogue. Constraints: one continuous shot, natural colour, no on-screen text.

TweakSay "one continuous shot" explicitly. Omni reasons about what should happen next, which means it will happily invent a cut if you leave the shot count open.

Gemini Omni Flash / All use cases

Image input — animate a still you already have

Prompt
Input role: use the attached image as the exact first frame — keep the subject, framing, and colour grade unchanged. Goal: bring the scene to life for 8 seconds without changing what is in it. Motion: **the steam rises from the cup, the light shifts slightly as a cloud passes**, camera holds locked-off. Audio: quiet cafe room tone, a spoon against ceramic once. Constraints: do not add people, do not move the camera, do not re-frame.

TweakThe "input role" line is the part people skip. Without it Omni treats the image as loose inspiration rather than as frame one.

Gemini Omni Flash / All use cases

Voice reference — match a specific delivery

Prompt
Input role: use the attached audio only as a voice reference for the speaker''s tone and pacing — do not reuse the words. Goal: an 8-second piece to camera. Scene: a **woman in a linen shirt** sits in a **sunlit home office**, speaking directly to the lens. She says: "**This took me two years to figure out.**" Motion: static medium close-up, shallow depth of field. Audio: her voice matched to the reference, quiet room tone underneath, no music.

TweakVoice references are the audio input type available first. State plainly that the words come from your script and the reference supplies only delivery.

Gemini Omni Flash / All use cases

Video input — restyle an existing clip

Prompt
Input role: use the attached video as the motion and timing reference — keep the camera move and the subject''s action exactly as they are. Goal: restyle the footage. Scene: re-render the same action as **a 1970s 16mm film look — warm grain, slight gate weave, softer contrast**. Audio: keep the original ambience, add no music. Constraints: do not change framing, do not change the length, do not add or remove anything from the scene.

TweakSeparating what to keep from what to change is the whole job. List the keeps first — Omni holds them better when they lead.

Gemini Omni Flash / All use cases

Conversational edit — the follow-up instruction

Prompt
Keep everything as it is. Change only the lighting: make it **late afternoon instead of midday**, warmer key from frame left, longer shadows. Do not change the subject, the camera move, the audio, or the length.

TweakThis is a second-turn prompt, not a standalone one. Omni's editing builds on the previous instruction, so name only the delta and explicitly fence everything else.

Gemini Omni Flash / All use cases

Product video — hero shot with native audio

Prompt
Goal: an 8-second product hero shot for a landing page. Scene: a **brushed-aluminium water bottle** stands centred on a **seamless charcoal surface**, condensation beading on the metal. Motion: slow 180-degree orbit, camera at product height, shallow depth of field. Audio: a low room tone and one soft metallic ring as the camera settles. Constraints: one continuous shot, no on-screen text, no music, no hands in frame.

TweakSwap the bold product and surface. Naming a single audio event tied to a moment gives Omni something to synchronise against.

Gemini Omni Flash / All use cases

Physics-led shot — use what the model is good at

Prompt
Goal: an 8-second single shot built around a physical event. Scene: a **ceramic mug of black coffee** on a **wooden table**; a hand nudges it and it **tips, and the liquid arcs out and spreads across the grain**. Motion: locked-off macro, shallow depth of field, no camera movement. Audio: the ceramic knock, the liquid hitting wood, no music. Constraints: one continuous shot, real-time speed, no slow motion.

TweakFluid, weight, and momentum are what Omni's physics understanding is built for. Describe the event, not the aesthetic, and let the model resolve the motion.

Gemini Omni Flash / All use cases

Real-world grounding — lean on Gemini's knowledge

Prompt
Goal: a 10-second historically grounded establishing shot. Scene: a **working printing house in 1840s London** — compositors at type cases, an iron hand-press in the foreground, gaslight and window light mixed. Motion: slow push-in past the type cases toward the press. Audio: the press thumping, type being set, muffled street noise outside. Constraints: period-accurate equipment and dress, one continuous shot, no on-screen text, no music.

TweakOmni is grounded in Gemini's knowledge of history and culture, so a named period and place outperforms a generic "old-timey" description.

Gemini Omni Flash / All use cases

Vertical social clip

Prompt
Goal: a 9:16 vertical clip for Shorts, 8 seconds, one continuous shot. Scene: a **street food vendor** folding a **crepe on a hot plate** in a **night market**, steam and neon behind. Motion: handheld, slight natural sway, framed tight on the hands. Audio: the batter hissing, the spatula scraping, crowd murmur behind. No music, no dialogue. Constraints: vertical framing, no on-screen captions, do not cut away.

TweakState the aspect ratio in the goal line rather than at the end. Framing decisions made early hold better than ones appended last.

Gemini Omni Flash / All use cases

Dialogue — two speakers, one shot

Prompt
Goal: an 8-second two-hander, one continuous shot. Scene: **two colleagues at a kitchen counter in an office**, one leaning on the counter. The first says: "**You told them already?**" The second, not looking up: "**Someone had to.**" Motion: static medium two-shot, shallow depth of field. Audio: both lines clean and synced, office room tone underneath, no music. Constraints: no on-screen captions, do not cut between speakers.

TweakKeep the combined dialogue short. Two brief lines sync more reliably in a short clip than one long speech, and "do not cut between speakers" stops Omni from editing it into shot-reverse-shot.

Gemini Omni Flash / All use cases

Multi-input — image plus text direction together

Prompt
Input role: the attached image is the **product and its exact colour and finish**; do not restyle it. Goal: place that product into a new scene for 8 seconds. Scene: the product sits on a **rain-flecked window ledge** with **a blurred city street below**, overcast daylight. Motion: slow vertical crane down past the ledge, ending centred on the product. Audio: rain on glass, faint traffic, no music. Constraints: product geometry and colour unchanged, one continuous shot.

TweakOmni reasons across inputs together rather than stitching them in turn, so state what each input is FOR. Here the image is a fidelity constraint, not a first frame.

Gemini Omni Flash / All use cases

Negative-led prompt — when output keeps drifting

Prompt
Goal: 8 seconds, one continuous locked-off shot, nothing else. Scene: an **empty classroom in late afternoon**, dust in the light, chairs on desks. Motion: none — the camera does not move at all. Audio: room tone only. Constraints: no camera movement of any kind, no people, no music, no on-screen text, no cuts, no zoom, no colour grading beyond natural daylight.

TweakWhen a scene keeps drifting, restate the constraint in both the motion line and the constraints line. Saying "none" once is weaker than saying it twice in different words.

Gemini Omni Flash / All use cases

Cinematic reveal

Prompt
Goal: a 10-second cinematic reveal, one continuous shot. Scene: a **lone figure in a heavy coat** stands at the edge of a **frozen reservoir at blue hour**, back to camera; ahead, **a half-submerged church spire breaks the ice**. Motion: slow crane up and forward, the spire entering frame as the camera rises. Audio: wind across ice, the ice groaning once, a single low sustained note. Constraints: one continuous shot, natural desaturated colour, no on-screen text.

TweakGive the reveal a trigger — "the spire entering frame as the camera rises" ties the payoff to the camera move instead of leaving the timing open.

Gemini Omni Flash / All use cases

Loopable ambient clip

Prompt
Goal: an 8-second clip that loops cleanly — the last frame should match the first. Scene: **rain running down a window** with **warm interior light** behind it, out of focus. Motion: locked-off macro, no camera movement. Audio: steady rain, no dynamics, no music. Constraints: no beginning or ending event, no change in light level, no cuts, nothing that would make the loop point visible.

TweakLoops need the absence of events stated explicitly. Any single moment — a drop landing, a light flickering — makes the seam visible.

FAQ

What is Gemini Omni Flash?

Gemini Omni Flash is Google's multimodal generation model, announced at I/O 2026. It takes images, audio, video, and text as input and generates video with sound. Google describes it as creating "anything from any input — starting with video," with image and audio output planned later. It is a separate line from Veo, not a Veo release.

Is Gemini Omni Flash the same as Veo 4?

No. There is no Veo 4. At I/O 2026 Google's next-generation video model was announced as Gemini Omni Flash, and Veo 3.1 remains the current Veo line. Pages advertising "Veo 4 prompts" are describing a model that does not exist.

How do I access Gemini Omni Flash?

It rolled out to Google AI Plus, Pro, and Ultra subscribers through the Gemini app and Google Flow, and at no cost to users on YouTube Shorts and the YouTube Create app. Developer and enterprise access came through the API after launch. The free YouTube route is the cheapest way to try these prompts.

How is prompting Gemini Omni different from prompting Veo?

Two differences matter. First, Omni accepts image, audio, and video inputs, so a prompt has to say what each input is FOR — first frame, style reference, voice reference, or fidelity constraint. Second, Omni edits conversationally: follow-up instructions build on the previous result rather than replacing it, so a second-turn prompt should name only the change and fence everything else.

Does Gemini Omni generate audio?

Yes. It produces video with sound, so audio direction belongs in the prompt the same way it does for Veo — name the ambience, tie sound effects to visible actions, and say "no music" explicitly if you do not want a music bed.

What is Gemini Omni Flash bad at?

Google's own model card names three: holding complete consistency across a sequence of edits, scenes with complex motion, and rendering accurate text. Changing a person's speech during editing is also deliberately restricted. Frame lettering small or angled rather than fighting the text rendering.

Are these prompts verified?

No. Every prompt on this page is labelled untested. They are written to the model's documented behaviour and to the structure the format needs, but they have not been run against a real Gemini Omni output and no public evidence video is attached. Treat them as starting structures.