Prompt Workbench
Veo Prompt Builder
To write a good Veo prompt, name the subject, the action, the camera, the light and the sound, in that order, and write it for the clip Veo can actually make: 4, 6 or 8 seconds, in 16:9 or 9:16. This builder turns a short brief into that structure, including the audio direction and a negative line, and gives you the result as paste-ready text or JSON.
Free with Google sign-in
AI Brief Assist
Paste a rough idea, product note, or messy offer and let AI fill the brief fields before you tune the final prompt.
Free while we test it. Buy me a coffee if it saves you time.
Veo rarely fails because your idea is weak. It fails because a line like “a woman drinking coffee, cinematic” leaves the model to invent the shot, the lens, the light, the pacing and the sound, and every field it invents can change on the next generation. The builder above closes those gaps before you spend a credit. It takes a brief and assembles the parts a video model reads: subject, action, setting, camera, lighting and audio.
It runs in your browser. Nothing is uploaded, there is no account, and there is no generation cost, because the tool writes prompts rather than video. You take the output to Veo, or to whichever model you run.
What is different about prompting Veo
Most of what makes a Veo prompt work is the same as for any text-to-video model. A few things are specific to Veo, and most of them are limits you should plan around before you write a word.
Clip length is 4, 6 or 8 seconds. Those are the only durations Google’s Veo 3.1 documentation lists. Eight seconds is also the only length that works with 1080p, 4K, reference images and extension, so if you want the sharpest output or you want to feed Veo reference images, you are writing for exactly eight seconds. The builder offers only these three lengths, and it writes the length into the prompt as, for example, “8-second”.
Aspect ratio is 16:9 or 9:16. Nothing else. 16:9 is the default. If a brief says “square for the feed”, decide now whether you are cropping later or switching to 9:16, because Veo will not generate 1:1.
Audio is always on. Veo 3.1 generates sound with the picture. Google’s prompting notes give three cues, and they behave differently:
- Dialogue: put the spoken words in quotation marks.
- Sound effects: describe them explicitly (“tires screeching”, “the latch clicks shut”).
- Ambient noise: describe the soundscape of the place.
Our advice, from working with these cues rather than from a Google rule: keep a spoken line short enough to say comfortably in the clip (roughly a dozen words for an 8-second shot), put the tone outside the quotation marks, and name the ambience so the model does not pick one for you. This is guidance we have not measured, which is why untested prompts on this site carry an explicit label.
What Google tells you to put in a prompt. The Gemini API documentation lists seven elements: subject, action, style, camera position and motion, composition, focus and lens effects, and ambiance (colour and light). The first three are the core; the rest are optional but are exactly where run-to-run variation hides. The builder has a field for every one of them.
Negatives are written as plain text. Google’s Veo documentation describes no separate negative-prompt field, so the builder puts the line inside the prose as “Negative prompt: …”. Phrase it as a description of what is absent (“no on-screen captions, no background music, no zoom”). Whether any given negative holds is something you have to test per shot; we do not claim a guarantee.
Reference images and first/last frame. Veo 3.1 and 3.1 Fast accept up to three reference images of a person, character or product, and Veo 3.1 Lite does not. First-frame and last-frame control is available on all three variants. When you animate from images, describe only what moves, how the camera behaves and what is heard; the image already shows the rest. The Veo image-to-video prompts page works through that pattern.
A before and after, in the builder’s structure
What most people type:
A woman showing a skincare serum, UGC style, vertical, 8 seconds.
What the builder assembles from a short brief of product, hook, scene and proof:
Create an 8-second 9:16 TikTok video ad for a vitamin C serum, optimized
for Veo 3.1. Objective: win first purchases from skincare buyers in their
late twenties. The opening line is: "Three weeks. That's it." Scene: a sunlit
bedroom, a woman in her late twenties sitting on the edge of the bed and
holding a frosted-glass serum bottle toward the lens. Camera and look:
handheld selfie framing with slight lens breathing, natural window light
from frame left, soft shadows. Show the proof or benefit: she tilts her
cheek into the light so the skin is visible, through visible action, not
text overlays. Audio: she says the line to camera, with room tone and a
faint street hum underneath. Brand constraints: label faces the camera,
no health claims. Final CTA beat: she holds the bottle at the lens for the
last second. Negative prompt: no on-screen captions, no background music,
no zoom, no watermark, no warped label.
This is a prompt structure, not a tested result. No Veo output has been verified for this exact prompt, and we do not show a video for it. It illustrates how the builder arranges a brief.
Read what each clause removes. The format line fixes the crop and the length. The named light direction fixes the mood across takes. The quoted line and the audio sentence fix the lip-sync target and the sound bed. The closing negatives aim at the things that tend to appear unasked on short social clips: captions, a music bed, a drifting zoom. Whether each one holds on Veo is the part you verify.
One consequence of the 8-second limit is pacing. The builder’s four-beat plan (hook, product, proof, call to action) cannot give four equal beats in eight seconds. A split that usually reads cleanly is a hook in the first two seconds, product and proof in the middle, and a still hero frame for the last second or two. That split is our pacing suggestion, not a Veo rule. If a story genuinely needs more time, generate a second clip and extend rather than cramming more beats in.
Text output or JSON output
The builder produces the same prompt in two shapes.
- Text is the paste-ready prompt for the Veo app, Gemini or any chat-style interface: one block of prose, with the most important elements first. Use it for one-off shots.
- JSON is the same prompt split into named fields. Use it for a batch where shots must match: you can change the subject across ten variants while the camera, lighting and negative fields stay byte-identical. It also diffs cleanly in version control once a prompt becomes an asset you maintain.
Neither is better. JSON is not a hidden mode that unlocks quality; it is the same instruction with structure you can hold steady. The gain is repeatability, not fidelity. The JSON prompt format guide covers when it pays off.
Veo compared with the other models in this builder
The builder writes for several models, and the limits differ enough to change the brief. Figures are from each vendor’s own documentation as of 2026-10-03.
| Veo 3.1 | Kling 3.0 | Seedance 2.5 (2.0) | Grok Imagine 1.5 | |
|---|---|---|---|---|
| Clip length | 4, 6 or 8 s | 3 to 15 s | 4 to 30 s (4 to 15 s) | 1 to 15 s |
| Aspect ratios | 16:9, 9:16 | 16:9, 9:16, 1:1 | six, including 16:9, 9:16 and 1:1 | seven, including 16:9, 9:16 and 1:1 |
| Audio | native, always on | native | native | native, on by default |
The practical difference: Veo is the most constrained on length and shape, so a Veo brief has to be the tightest. If you need a 12-second clip or a square crop, that is a reason to look at Kling, Seedance or Grok Imagine rather than to fight Veo’s limits. If you want Google’s newer conversational-editing model, see the Gemini Omni prompts page. The structure you build here carries across all of them; what you adjust is length, aspect ratio, audio and negatives.
Which Veo variants this works with
The builder is built for the current Veo line: Veo 3.1, Veo 3.1 Fast and Veo 3.1 Lite. Google’s changelog lists those three, and Google’s Veo page names Veo 3.1 as its latest model. There is no Veo 4. Because the output is plain prose rather than proprietary syntax, it transfers to most other video models with light editing. Audio direction and negative prompts are the two parts most likely to need adjusting when you move.
Pick the variant for the job. Google describes 3.1 Lite as its most cost-efficient Veo model, designed for rapid iteration, so Lite and Fast are sensible places to test whether a subject, camera move and line of dialogue work at all. Use full Veo 3.1 or 3.1 Fast when you need reference images or extension, and full 3.1 or Fast for 4K. Lite supports none of reference images, extension or 4K.
Where to run the prompt
Official first. Google says Veo 3.1 is available in the Gemini app, Flow, YouTube Shorts and Google Vids, and to developers through the Gemini API, Google AI Studio and Vertex AI. The Gemini API documentation is where the limits described above are specified. Third-party platforms that run Veo 3.1 include PixVerse, which lists it alongside its own and other models, and Pollo AI, whose API documents the same 4, 6 and 8 second options, 16:9 and 9:16 framing, text-to-video and image-to-video. Which variants and settings you get depends on the platform and your plan, so check what your account offers. The prompt is the same wherever you run it.
Where to go next
If you would rather start from something already tried than from a blank brief, the Veo prompt library collects structures by use case, each labelled with how it was verified. The closest starting points to this builder are the UGC-style prompts, the product video prompts and the dialogue and audio prompts.
For the reasoning behind the structure, the guides cover the prompt formula, camera movement language and negative prompts for troubleshooting. The Veo 3.1 prompting guide puts them together in one page.
FAQ
How long can a Veo video be?
A single Veo 3.1 generation is 4, 6 or 8 seconds. Google requires 8 seconds for 1080p, 4K, reference images and extension. To go past 8 seconds you extend a Veo-generated clip by 7 seconds at a time, up to 20 times, at 720p; the Lite variant cannot extend.
Which aspect ratios does Veo support?
Veo 3.1 supports 16:9 (the default) and 9:16 only. Square 1:1 and other ratios are not available, so the builder only offers those two.
Is there a Veo 4 prompt builder?
No, because there is no Veo 4. Google lists Veo 3.1, Veo 3.1 Fast and Veo 3.1 Lite as its Veo models, and its newer video model is a separate line called Gemini Omni. This builder targets the Veo 3.1 family and is not tied to a version number, so it stays current when Google ships a new Veo.
Does Veo generate audio, and how do I direct it?
Yes. Audio is always on for Veo 3.1. Google documents three cues: put dialogue in quotation marks, describe sound effects explicitly, and describe the ambient soundscape. The builder writes all three into the prompt so Veo does not invent the soundtrack for you.
Sources and date checked
Veo limits on this page were checked against the sources below on . Model limits change; if a source disagrees with this page, trust the source.
- Gemini API: Generate videos with Veo 3.1
- Gemini API: Release notes
- Google DeepMind: Veo
- Google: Veo 3.1 Ingredients to Video (where Veo 3.1 is available)
- PixVerse: Kling O3 and Kling 3.0 (states Veo 3.1 is in the same workspace)
- Pollo AI API docs: Veo 3.1
- Kling AI: Text to Video (comparison table)
- CapCut newsroom: Dreamina Seedance 2.0 (comparison table)
- xAI API docs: Video generation (comparison table)