Prompt Workbench

Kling Prompt Builder

To write a Kling prompt, describe the subject, what it does, and the scene in plain, direct sentences, then add camera, lighting and mood. That is Kling's own framework, and you size the clip to what Kling 3.0 makes: 3 to 15 seconds in 16:9, 9:16 or 1:1, with sound generated alongside the picture. This builder arranges your brief in that order and writes the audio direction for you.

Free with Google sign-in

AI Brief Assist

Paste a rough idea, product note, or messy offer and let AI fill the brief fields before you tune the final prompt.

Login required

Free while we test it. Buy me a coffee if it saves you time.

01

Campaign

Mode

Kling VIDEO 3.0: 3 to 15 seconds, native audio with lip-synced speech, and multi-shot generation (up to six shots in one pass).

02

Offer

03

Video structure

04

Sound and guardrails

Kling does not fail because your idea is thin. It fails when a prompt like “a chef cooking, cinematic” leaves it to decide who the chef is, what the hands do, where the camera sits and what you hear. The builder above fills those gaps from a short brief, so the prompt you paste into Kling says the same thing every time you run it.

It runs in your browser. Nothing is uploaded, there is no account, and it costs nothing, because it writes prompts rather than video.

How prompting Kling differs

Kling publishes a prompt framework, and it is simple. Kling’s own prompt guide lists the parts of a good prompt as subject, visible action, scene, camera language, and lighting and mood. Its older text-to-video guide gives the same skeleton as subject, subject movement, scene, then camera language, lighting and atmosphere in brackets, meaning optional. The builder follows that order: who or what, what it does, where, then how it is filmed.

Keep the language plain. That older guide, which Kling wrote for an earlier model version, advises simple words and sentence structures and avoiding overly complex language. It also warns that numbers are unreliable and that complex physical movement, like a ball’s trajectory, is hard. Treat those as habits to keep until you have tested otherwise on 3.0: one clear action per subject, no counting (“three birds”), no choreographed physics.

Length is flexible, from 3 to 15 seconds. Kling 3.0 generates clips of 3 to 15 seconds, and Kling frames the longer end as room for “meaningful arcs”. That is nearly double Veo’s longest clip, so a Kling brief can hold a setup and a payoff where a Veo brief cannot. Longer is not automatically better: a single action needs only a few seconds, and the builder defaults to 5.

Square video is on the menu. Kling lists 16:9, 9:16 and 1:1. Resolution depends on your plan: Kling’s page says 720p is available on the free tier, with 1080p and 4K on higher plans.

Sound is generated, so write it. Kling 3.0 produces audio with the video: dialogue, sound effects and ambience, with lip-synced speech in English, Chinese, Japanese, Korean and Spanish and regional accents such as British and Indian English. Kling’s 3.0 guide advises pairing a character’s name with their line and stating the emotional tone, so the model knows who is speaking. The builder writes an audio sentence into every prompt for that reason; delete it if you want silence or plan to score the clip later.

Multi-shot is a real feature. Kling describes an “AI Director” that can produce a sequence of up to six distinct shots in one pass, and a custom multi-shot mode that gives you shot-level control. If you use it, write each shot as its own short line with one camera setup and one action, in the order they should appear. That structure is our advice, not a Kling-published template, and we have not tested it.

Negatives go inside the prompt, and short. Many third-party guides describe a separate “negative prompt” box for Kling. Kling’s current API reference for Kling 3.0 and 3.0 Omni has no such parameter: the prompt field (up to 3,072 characters, 2,500 or fewer recommended) is documented as able to include positive and negative descriptions. The separate negative_prompt field belonged to the legacy API. We have not inspected the signed-in app’s create form, so if your interface shows a negative box, use it for the same short list. The builder writes negatives inline as a closing “Negative prompt: …” sentence. Keep that list to a few specific things, and say what you want positively everywhere else (“static background”, “clean hands”, “natural skin”).

Image-to-video follows the same rule as everywhere. Kling’s image-to-video guide, written for an earlier version, says the image already provides the scene, so the prompt only needs the subjects and the movement you want. Describe motion, not the picture. Our Veo image-to-video prompts page works through the same principle with examples.

Kling 3.0 and 3.0 Omni. Omni is built for reference control: you call reference elements into the prompt with an @ symbol, for example @Grace or @Element1. Kling’s Omni guide allows up to 7 images or elements without a video input and up to 4 with one. Use plain 3.0 for text-to-video and single-image animation, and Omni when the same character or product has to hold across shots.

What about Kling 4.0?

Kling 4.0 is not generally available as of 2026-10-04, and Kling’s app still labels it early access (“Coming Soon”, with early access to Kling 4.0 Flash). Kling’s own announcement says early access began rolling out to a limited group of users on September 28, 2026 and that the full model is scheduled to launch in October 2026. The announced specifications include 3 to 30 second clips, 4K output, up to 10 keyframes, up to 15 combined references and prompts up to 8,000 tokens, with 10-bit HDR marked “coming soon”. Those are announced capabilities, not something you can count on in an ordinary account, so this page writes for 3.0. When 4.0 launches, the limits above will change and this guide and the builder’s options will be updated; the structure itself will not.

A before and after, in the builder’s structure

What most people type:

A man making coffee in a kitchen, nice lighting, 10 seconds.

What the builder assembles from a brief of product, hook, scene and proof (9:16, 10 seconds):

Create a 10-second 9:16 Instagram Reel video ad for a hand-pour coffee
kettle, optimized for Kling 3.0. Objective: win first purchases from home
coffee enthusiasts. The opening line is: "Slow coffee is a ritual." Scene:
a morning kitchen counter, a man in a grey apron pouring from a matte
black gooseneck kettle onto a dripper. Camera and look: slow push-in
from a medium shot to a close-up of the spout, soft window light from
frame left, warm and unhurried. Show the proof or benefit: a thin,
steady stream and a visible bloom on the grounds, through visible
action, not text overlays. Audio: the man says the line quietly to
camera; water pouring and a faint kettle tick, room tone underneath.
Brand constraints: kettle stays matte black, logo not required in
frame. Final CTA beat: the camera settles on the kettle on the counter
beside the finished cup. Negative prompt: on-screen text, watermark,
warped hands, extra fingers, flickering light.

This is a prompt structure, not a tested result. No Kling output has been verified for this prompt, and the negative list is a plain example, not a proven set. It shows how the builder arranges a brief.

Notice what changed. The format line fixes length and crop. The scene names the person, the object and the surface instead of leaving them open. The camera line is a single move, and the lighting names a direction. The audio sentence says who speaks and what the room sounds like. Each is a choice Kling would otherwise make for you, and may make differently next time.

Kling compared with Veo

Both write sound with the picture and both reward a prompt built in a fixed order. Where they differ is what you can ask for.

Kling 3.0Veo 3.1
Clip length3 to 15 seconds4, 6 or 8 seconds
Aspect ratios16:9, 9:16, 1:116:9, 9:16 only
Multi-shot in one passup to six shotsnot documented in Google’s Veo guide
Audionative, lip-synced in five languagesnative, always on
Negativeswritten inside the prompt (current API has no separate field)no negative field documented; written in the prompt

In practice a Kling brief can carry more: a longer beat, a second shot, a square crop. A Veo brief has to be tighter because eight seconds is the ceiling. If your idea fits in eight seconds in a 16:9 or 9:16 frame, either can work, and trying the same brief on both is a fair way to choose. The Veo Prompt Builder uses the same structure with Veo’s limits, and the Veo prompt library shows the Veo side by use case.

Where to run the prompt

Official first: Kling AI (kling.ai) is Kling’s own platform and the place the 3.0 features described above are documented; plan tiers decide the resolution you get, with 720p on the free tier and 1080p and 4K on higher plans, per Kling’s page. Other platforms that run Kling: PixVerse lists Kling 3.0 and Kling O3 (Kling’s Omni model) with 3 to 15 second video, and its page says Kling video generation there requires a Pro plan or higher; Pollo AI documents Kling 3.0 Omni with 3 to 15 second clips and the same three aspect ratios. Kling 4.0 is still early access on Kling’s own platform only (checked 2026-10-04), as described above. The prompt is the same wherever you run it. If you want to try the same brief on a different model afterwards, the Seedance builder and the Grok Imagine builder take the same inputs.

FAQ

How long can a Kling video be?

Kling 3.0 generates 3 to 15 seconds in one clip, in 16:9, 9:16 or 1:1. Kling 4.0 has been announced for 3 to 30 seconds, but it is still early access as of October 4, 2026, with the full model scheduled for October 2026, so this builder targets the 3.0 limits.

Does Kling support negative prompts?

Yes, inside the prompt. Kling's current API reference for Kling 3.0 and 3.0 Omni has no separate negative prompt parameter; it says the prompt can include positive and negative descriptions. The separate negative_prompt field belonged to the legacy API. This builder therefore adds a short "Negative prompt: ..." sentence at the end of the prompt. Keep it to a few specific things to avoid, and describe everything you do want in positive terms.

Does Kling generate audio?

Yes. Kling 3.0 produces native audio with the video, including lip-synced speech in English, Chinese, Japanese, Korean and Spanish plus regional accents. Direct it in the prompt by naming who speaks and what they say, and by describing the sound you expect.

Is Kling 4.0 available yet?

Not generally. Kling's own blog says early access began on September 28, 2026 for a limited group of users and that the full Kling 4.0 model is scheduled to launch in October 2026. As of October 4, 2026, Kling's app still shows a "Coming Soon" banner with early access to Kling 4.0 Flash. 4K output is described as supported, while 10-bit HDR is labelled coming soon. Until it ships, write for Kling 3.0.

Sources and date checked

Kling limits on this page were checked against the sources below on . Model limits change; if a source disagrees with this page, trust the source.