Prompt Workbench
Grok Imagine Prompt Builder
To write a Grok Imagine prompt, give it a starting image and describe only what should move, how the camera behaves and what is heard. xAI's own example is a single sentence of motion and camera, and it never re-describes the picture. Starting from text works too, but then Grok builds a first frame from your words before it animates, so the opening composition has to be in the prompt.
Free with Google sign-in
AI Brief Assist
Paste a rough idea, product note, or messy offer and let AI fill the brief fields before you tune the final prompt.
Free while we test it. Buy me a coffee if it saves you time.
Grok Imagine is the model on this site where the image does most of the work. Give it a good still, and the prompt shrinks to a few clauses of motion, camera and sound. That is a different habit from writing a text-to-video prompt, where the words have to build the whole world, and it is why this builder has an image-to-video mode rather than only a brief form.
It runs in your browser. Nothing is uploaded, there is no account, and it costs nothing, because it writes prompts rather than video. Your image goes to Grok, not to us.
Two ways to start, one rule for each
Grok Imagine Video 1.5 is xAI’s current video model; xAI announced it on June 16, 2026, and its API id is grok-imagine-video-1.5. It has two main modes, and the right prompt is different for each.
Image-to-video: write only the change. xAI’s launch post describes the workflow as giving it a starting image and describing the motion. Its example prompt is one sentence: “Slow cinematic push-in as embers drift across the battlefield and the helmet’s crest stirs in the wind.” It names a camera move and two things that move. It does not describe the helmet, the battlefield or the lighting, because the image already shows them. PixVerse, which also runs Grok Imagine 1.5, gives the same advice in its own guide: focus the prompt on motion instead of recreating the scene.
Text-to-video: write the first frame, then the motion. xAI’s API docs explain that text-to-video generates a first frame from your prompt and then animates it, in a single request, without showing you the intermediate image. The consequence is practical. The first thing you describe sets the opening composition, so put subject, framing and light up front, then the motion. A prompt that starts with action and never says what the first frame looks like gives the model a free hand with the one frame everything else grows from. That reading of the docs is our inference, not an xAI-published tip.
How prompting Grok Imagine differs
Image-to-video habits that work across models. These are our advice, adapted from the image-to-video pattern in our Veo image-to-video prompts page, and they are untested on Grok specifically:
- Describe motion, not the scene. “Steam curls up from the mug” rather than a new description of the mug.
- Use weighted verbs. Sway, ripple, drift and push give motion some weight; “moves” and “animates” do not.
- Name the camera move. If you do not, you are leaving it to the model. “Slow push-in”, “locked-off” and “gentle handheld sway” are all instructions.
- Keep it small. One or two clear actions beat a busy performance, particularly with faces and products.
- Name objects that are in the image. If the coffee cup is on the table, say “the coffee cup”.
Sound is generated in the same pass. xAI says sound effects, ambience and dialogue are generated together with the video and “land on the action”, and that speech is clearer and better synced than in the previous version. So direct the audio in the prompt: name one sound tied to a visible action, name the ambience, and put any spoken line in quotation marks with the tone stated outside them. xAI has not published a dialogue syntax that we found, so the quotation-mark convention is borrowed from other models and worth testing.
Length is up to you, from 1 to 15 seconds. The API accepts 1 to 15 seconds. A one-action clip rarely needs more than 4 to 6 seconds, and the builder defaults to 6, which matches xAI’s own examples. More detailed scenes also take longer to process, according to the API docs.
Aspect ratio: leave it alone in image-to-video unless you mean it. The API supports seven ratios: 1:1, 16:9, 9:16, 4:3, 3:4, 3:2 and 2:3. For image-to-video the output follows the input image’s ratio by default, and the docs warn that overriding it with the aspect_ratio parameter stretches the image. If your picture is 4:5 and you want 9:16, crop or extend the picture first and then animate it.
Resolution. The API offers 480p by default, with 720p and 1080p, and 1080p is limited to the 1.5 model for text-to-video and image-to-video. Reference-to-video is capped at 720p. The apps and partner platforms may expose fewer options.
References. xAI’s July 31 post describes reference images that “lock one thing in place”, such as a face, a product or a location, with up to seven references per generation, plus a voice reference so the same face and voice hold across scenes. At launch, references were rolling out first to SuperGrok Heavy and SuperGrok Plus subscribers in the US, then to all tiers. If you use them, say what each one is for in the prompt, which is our advice.
Negatives: we have no documented guidance. No xAI page we found mentions negative prompts, and the video API has no negative parameter. State what you want in positive terms (“static background”, “clean hands”) and keep anything you want to avoid short. For Grok Imagine the builder leaves the negative prompt out of its output and shows a tip to phrase what you want positively instead.
A before and after, in the builder’s structure
What most people type, with a photo of an earbuds case attached:
Make my earbuds case photo look cool and animated.
Below is what the builder assembles for its “Product demo” preset. You can reproduce both versions at /tool/grok-imagine/?preset=product-video&mode=i2v and by switching the mode to Text-to-video. First, image-to-video, where the image stays fixed and the words go on motion, camera, proof, audio and the closing beat:
Animate the uploaded image into a 6-second 16:9 landing page hero video
ad for a matte-black wireless earbuds case, optimized for Grok Imagine
Video 1.5. The clip starts from the uploaded image; keep its subject,
setting and look unchanged. Objective: sell a product to commuters who
want compact earbuds with strong battery life. The opening line is:
"Small case, all-day sound". Motion and camera: smooth orbit around the
product. Show the proof or benefit: the lid opens smoothly, LED battery
indicator glows, earbuds snap in magnetically through visible action,
not text overlays. Audio: subtle upbeat music, soft case click, light
magnetic snap, no voiceover. Brand constraints: premium, minimal,
precise, no fake interface labels. Final CTA beat: Show the final
product centered and ready to buy.
Switch to text-to-video and the builder adds the scene, because the model has no picture to start from. It also moves the lighting into the camera and look line:
Create a 6-second 16:9 landing page hero video ad for a matte-black
wireless earbuds case, optimized for Grok Imagine Video 1.5. Objective:
sell a product to commuters who want compact earbuds with strong battery
life. The opening line is: "Small case, all-day sound". Scene: minimal
studio tabletop setup with hands-only product demonstrator. Camera and
look: smooth orbit around the product, high-key clean studio lighting.
Show the proof or benefit: the lid opens smoothly, LED battery indicator
glows, earbuds snap in magnetically through visible action, not text
overlays. Audio: subtle upbeat music, soft case click, light magnetic
snap, no voiceover. Brand constraints: premium, minimal, precise, no
fake interface labels. Final CTA beat: Show the final product centered
and ready to buy.
These are prompt structures, not tested results. No Grok Imagine output has been verified for either prompt, and we show no video for them. They illustrate how the builder arranges the same brief in each mode.
What the image-to-video version does: it tells the model the clip starts from your image and that the subject, setting and look must stay unchanged, then spends every remaining word on movement, camera, the proof shown on screen, sound and the final beat. That is the logic of xAI’s own one-sentence example, expanded just enough to cover audio and a closing action. What the text-to-video version does: it keeps the same brief but adds a Scene line and a lighting note, so the first frame the model builds is described rather than left to chance. Neither version ends with a negative-prompt line, because for Grok Imagine the builder does not export one.
Grok Imagine compared with Veo
| Grok Imagine Video 1.5 | Veo 3.1 | |
|---|---|---|
| Clip length | 1 to 15 seconds | 4, 6 or 8 seconds |
| Aspect ratios | seven, including 16:9, 9:16 and 1:1 | 16:9, 9:16 only |
| Starting from an image | one starting image; up to seven references | image-to-video, first and last frame, up to three reference images (not on Lite) |
| Audio | native, on by default | native, always on |
| Top resolution | 1080p (API) | 4K, on 8-second clips only |
| Extending a clip | extension endpoint in the API | 7 seconds at a time, up to 20 times, 720p only |
Choose by what you have. Grok Imagine is flexible on length and shape and happy with a single photo and a short prompt, which suits quick social clips. Veo gives you the first-and-last-frame control and a documented extension route, and its limits on length and shape are tighter. For an image you want to animate precisely, our Veo image-to-video prompts show the same describe-the-motion discipline with examples, and the Veo Prompt Builder applies it to Veo’s limits. If you would like to compare other models on the same brief, there are builders for Kling and Seedance too.
Where to run the prompt
Official first: Grok Imagine runs on grok.com/imagine and in xAI’s iOS and Android apps, where xAI says text-to-video and 1080p are generally available, and through the xAI API as grok-imagine-video-1.5. Other platforms: PixVerse offers Grok Imagine 1.5, but only in Image-to-Video mode, for Pro members and above, at 480p or 720p, and its audio is on by default and cannot be turned off, so use the image-to-video prompts above there rather than the text-to-video one. Runway lists grok_imagine_1_5 in its API, accepting text or image input. Options, resolution and plan requirements differ by platform, so check what your account shows. The prompt is the same wherever you run it.
FAQ
Does Grok Imagine do text-to-video, or only image-to-video?
Both. xAI's API docs list text-to-video and image-to-video for Grok Imagine Video 1.5, and xAI says text-to-video and 1080p are generally available on grok.com/imagine and its iOS and Android apps. Some partner apps only expose image-to-video: on PixVerse, Grok Imagine 1.5 is available in Image-to-Video mode only.
How long can a Grok Imagine video be?
The xAI API accepts 1 to 15 seconds per clip, at 480p (the default), 720p or 1080p. The app menus may show a shorter list, so check the options on the surface you use. PixVerse offers 1 to 15 seconds at 480p or 720p.
Does Grok Imagine generate sound?
Yes. xAI says sound effects, ambience and dialogue are generated in the same pass as the picture, and the API turns audio on by default. You can ask for silent output through the API, but on PixVerse the audio is on and cannot be turned off.
What should I write in a Grok Imagine image-to-video prompt?
Only the change: what moves, how the camera moves, and what we hear. The image already shows the subject, setting and style, so describing them again gives the model room to drift away from your picture. Name objects that are already in the frame when you want them to move.
Sources and date checked
Grok Imagine limits on this page were checked against the sources below on . Model limits change; if a source disagrees with this page, trust the source.
- xAI: Grok Imagine Video 1.5 (2026-06-16)
- xAI: Imagine Video 1.5 with References (2026-07-31)
- xAI API docs: Video generation
- xAI API docs: Imagine overview
- PixVerse: Grok Imagine 1.5 on PixVerse (runs this model, Image-to-Video only)
- Runway API: Available models (lists Grok Imagine Video 1.5)
- Gemini API: Generate videos with Veo 3.1 (comparison)