Writing Prompts for AI Video Models: Describe Causes, Not Effects (and Stop Stacking '8K Masterpiece')
There is one master rule for video prompts: describe what causes the picture, not the picture itself. Light: write the source's position, quality and colour temperature, never "light hits her forehead". Motion: contact point, weight, end state, never "fluid". Dialogue: who speaks and for how many seconds (about 2.5 words per second), never "emotional". "8K, masterpiece, volumetric" are CLIP-era leftovers; not one of Kling's, Google's or Alibaba's official guides asks you to stack them. They give you sentence slots.
Three official formulas, one lesson
Kling: Subject + Subject Movement + Scene + (Camera Language + Lighting + Atmosphere). Google, for Veo 3.1: [Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance]. Alibaba Model Studio, for Wan: Subject + Scene + Motion, plus Aesthetic Control + Stylisation in the advanced version. Every slot in all three is a fact you can state in one sentence. None has a slot called "quality words".
"masterpiece, best quality, 8K, ultra detailed" is an artefact of the Stable Diffusion 1.5 era: those tokens co-occurred with high-scoring images in training captions, so stacking them worked on a CLIP encoder. Modern video models read whole sentences through an LLM-style text encoder. The tokens add nearly nothing and crowd out slots that carry information; MiniMax H3 accepts 7,000 characters per text item, but attention is not spread evenly across them.
So the master rule is one line: describe the cause of the picture, never the result.
Lighting: write the source, not the landing spot
The most reproducible failure in production pipelines: the prompt says "warm yellow light brightens her forehead and nose bridge", and the output has a blob of yellow on the forehead and nose bridge, exactly where the prompt pointed, like a torch held to the face. The model executes a landing spot as "paint something here".
Write the cause instead: A warm-white pendant lamp casts soft light from the upper right of frame; cool window daylight fills from the left. Source position, quality (hard or soft) and colour temperature; where it lands is the model's arithmetic. Google's own Veo examples take this form: "harsh fluorescent overhead lights", "the green glow of the monochrome monitor". Sources and qualities, not landing spots.
Three companions. Never write the light/shadow boundary; falloff is a gradient, and naming a "boundary" draws a hard edge. Never use saturated colour names for temperature; 3200 K on skin is warm white, and "yellow" gets you yellow paint. Write volumetric light and god rays as a medium: "a layer of unsettled smoke hangs in the booth". The model knows what light does in smoke.
Motion: contact point, weight, ground relation, end state
"He angrily shoves the door open" leaves four things unspecified; the model guesses each.
- Contact point: palm at what height on the door, or a shoulder?
- Weight: solid oak or hollow core; does the body lean into it?
- Ground relation: feet inside or outside the threshold; does a step accompany the push?
- End state: door at what angle, body stopped where, hand still on the door or not?
End state matters most. A clip has an endpoint; leave it unstated and the model picks one, typically a return to the starting pose or a looping gesture. With several shots in one segment, each shot's end state is the next shot's start state; write it and the cuts stop jumping.
"Fluid", "natural", "powerful" are result words. Delete them.
Dialogue: speaker attribution and seconds
Say who speaks. Two people in frame and no attribution, and the lip-sync lands on the wrong face. Alibaba's Wan guide requires "Character says: 'line'"; Google's Veo example is "A woman says, 'We have to leave now.'"
Budget seconds by speech rate. English runs about 2.5 words per second, Mandarin about 5 characters: a 12-word line needs 4–5 seconds, plus roughly 0.75 s per terminal punctuation mark. Over-budget produces a very stable artefact: a 7-second shot with 4.5 seconds of dialogue and no closing action gets the line spoken twice to fill time. The fix is a closing beat at the end of every segment, a small action without displacement: setting down a cup, looking away. Under-budget, the back half of the line spills onto the next shot's face.
Wan's official voice formula: voice = spoken content + emotion + tone + pace + timbre + accent, example He says: "Study hard, make progress every day", relaxed tone, moderate pace, clear voice, American English. Sound effects = source material + action + ambient sound. Veo's convention: dialogue in quotes, sound effects prefixed SFX:, background prefixed Ambient noise:. Both agree: tone and pace are modifiers outside the quotes, never inside them.
Five things never to write
- Negations. "No fog", "no extra people" put fog and people into the conditioning. Google's guidance for Veo is to describe what you wish to exclude, for example "a desolate landscape with no buildings or roads" rather than "no man-made structures". Write the positive opposite.
- The wrong silence word. Wan's official control words are "No dialogue" (Chinese 无台词) and "No background music"; a paraphrase such as 无对白 reads as ordinary description and does nothing.
- Performance notes inside quotation marks. Wan uses quotes as its dialogue delimiter, so "(sneers)" inside them gets spoken. Put delivery outside the quotes, in the tone and pace slots.
- Camera moves on every shot. A locked-off camera is the default; write a push-in or a pan only with a narrative reason. Wan's control words are "static shot", "camera push-in", "camera pans left". Move on every shot and the cut looks like a commercial.
- Facing along the depth axis. "She turns and walks into the depth of the frame" buys three seconds of the back of her head. Write facing as frame-left or frame-right; express depth as a positional end state: "walks to the window and stops".
The sixth was covered above: SD trigger words. No official guide from Kling, Google or Alibaba lists masterpiece, 8K, subsurface scattering or volumetric lighting.
Before / after
| Goal | Effect (rewrite this) | Cause (write this) |
|---|---|---|
| Key light | Warm yellow light brightens her forehead and nose bridge | A warm-white pendant lamp casts soft light from the upper right of frame; cool window daylight fills from the left |
| Atmosphere | Volumetric lighting, god rays, cinematic | A layer of unsettled smoke hangs in the booth; the pendant's light passes through it |
| Motion | He angrily shoves the door open, powerful and fluid | His right palm slams the centre of the solid oak door; he leans into it, takes one step over the threshold and stops; the door rests at about 90° |
| Dialogue | She delivers an emotional line (7 s) | She says: "You knew about this all along." Low voice, slow pace; then her eyes drop to the rim of the cup (4.5 s) |
| Exclusion | No fog, no extra people | Clear, dry air; she is the only person in frame |
| Camera | Slow push-in, orbit, cinematic camera work | Static shot, medium framing, subject on the right third |
| Quality | masterpiece, 8K, ultra detailed | (delete; describe the look via film stock, focal length, depth of field) |
The left column states results, which the model can only paint; the right states causes, which it can compute.
Per-model quirks
| Model | Official formula | Dialogue and sound | Referring to references | Notes (official) |
|---|---|---|---|---|
| Kling | Subject + Subject Movement + Scene + (Camera Language + Lighting + Atmosphere) | — | by element (Element Library) | an element is 2–4 images of one subject from different angles |
| Veo 3.1 | [Cinematography] + [Subject] + [Action] + [Context] + [Style & Ambiance] | dialogue in quotes; SFX: / Ambient noise: prefixes | — | describe what to exclude instead of writing "no" / "don't" |
| Wan 2.7–3.0 (Alibaba) | Subject + Scene + Motion (+ Aesthetic Control + Stylisation) | voice = content + emotion + tone + pace + timbre + accent; "No dialogue", "No background music" | Image 1 / Video 1 | images + videos ≤5; "static shot", "single-shot" control words |
| Seedance 2.0 (Volcengine) | theme + visual style + frame detail and motion + format requirements (Volcengine article) | not detailed in the official article | Image 1 / Image 2 | multimodal reference and first/last frame are separate scenarios |
| MiniMax H3 | content array: text + references | reference audio ≤3 clips, ≤15 s total | by order | ≤7,000 characters per text item; references and first/last frame are mutually exclusive |
The overlap outweighs the differences: slot formulas, natural language, references by index or name. The differences cluster around sound: Wan and Veo have explicit dialogue syntax, the others lean on reference audio. Switching models means switching dialogue syntax and reference index; subject, motion and light are written the same everywhere.
A shot template that survives contact with the model
- Segment header: location; opening state (who is frame-left / frame-right, posture, what is in their hands); one lighting sentence (position + quality + colour temperature); static camera, or a move with a reason.
- One line per shot: seconds / framing / who is in frame doing what (contact point, weight, end state) / dialogue (speaker + quoted line + tone and pace outside the quotes) / lighting only where it differs from the header.
- Closing beat: one small action without displacement after the last line.
- Self-check: delete result adjectives (fluid, stunning, cinematic, 8K); convert negations to positives; recompute dialogue seconds at 2.5 words/s; no facing along the depth axis; references in this model's index format.
Key figures and sources
| Figure | Value | Source | Date |
|---|---|---|---|
| Wan official voice formula | voice = spoken content + emotion + tone + pace + timbre + accent; silence control words are "No dialogue" / "No background music" | help.aliyun.com | 2026-09-02 |
| Wan base prompt formula | Subject + Scene + Motion; advanced adds Aesthetic Control + Stylisation | help.aliyun.com | 2026-09-02 |
| Veo 3.1 exclusions | Google recommends describing what to exclude, e.g. "a desolate landscape with no buildings or roads", instead of "no man-made structures" | cloud.google.com | 2026-09-02 |
| Kling prompt formula | Subject + Subject Movement + Scene + (Camera Language + Lighting + Atmosphere) | kling.ai | 2026-09-02 |
| MiniMax H3 prompt limit | ≤7,000 characters per text item; reference audio ≤3 clips, ≤15 s total | platform.minimax.io | 2026-09-02 |
| Wan 2.7 reference index | references addressed as Image 1 / Video 1 in the prompt | help.aliyun.com | 2026-09-02 |
FAQ
Is a longer prompt a better prompt?
No. One factual sentence per slot. MiniMax H3 allows 7,000 characters per text item, but 60–120 words covering the five slots is usually enough; beyond that the additions are mostly adjectives.
Should I prompt in English or Chinese?
In the model's native language. Kling, Wan and Seedance publish their guides in Chinese; Google's Veo guide is in English. Do not swap native tone or physiognomy vocabulary for near-synonyms.
My model has a separate negative-prompt field. Use it?
Style-level exclusions can go there. Never put negations into the positive prompt itself; Google's Veo guidance says to describe the exclusion instead.
Why does the model speak the line twice?
Too many seconds and no closing action. A 7-second shot with 4.5 seconds of dialogue leaves 2.5 seconds with no exit, so the line repeats. Recompute at 2.5 words per second, or add a closing beat.
How do I write "cinematic"?
Write what produces it: focal length and depth of field (85 mm, shallow), light source and medium, colour discipline (low saturation, one accent colour), a static camera. The word itself carries no executable information.
How do I write several shots in one segment?
One line per shot with seconds and framing, naming only who is in frame, with each shot's end state as the next shot's start state. Wan also offers a "single-shot" control word for the opposite constraint.
Sources
- https://help.aliyun.com/zh/model-studio/text-to-video-prompt
- https://help.aliyun.com/zh/model-studio/wan-video-to-video-api-reference
- https://cloud.google.com/blog/products/ai-machine-learning/ultimate-prompting-guide-for-veo-3-1
- https://docs.cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide
- https://kling.ai/quickstart/klingai-video-o1-user-guide
- https://kling.ai/quickstart/klingai-element-library-3-user-guide
- https://platform.minimax.io/docs/api-reference/video-generation-v2-create
- https://www.volcengine.com/article/40840