HomeGuides › AI Short-Drama Storyboard Template: writing a 15-second segment a video model will actually execute

AI Short-Drama Storyboard Template: writing a 15-second segment a video model will actually execute

A storyboard a 2026 video model executes looks like this: a segment header (location, opening frame state, who stands screen-left/right) → 3–6 shots, each one line: "[Shot N] seconds: framing & camera / who is in frame and what they do / dialogue (speaker + voice hint) / lighting (source direction & colour temperature)" → a closing beat after the last line (a small non-locomotive action so the model does not fill leftover seconds by repeating dialogue). Three iron rules: write causes not results; list only who is in frame; size dialogue at ~2.5 words per second.

Segment template

[LOCATION] Uncle Chen's noodle shop — dining room (night, rain)
[OPENING STATE] Medium shot. Lin Wan sits on the long bench by the window, screen-left, both hands around a white bowl; Uncle Chen wipes his hands behind the stove, screen-right. Axis: Lin left, Chen right; all cameras this segment stay left of the axis.

[Shot 1] 4s: Medium, camera level with the stove. Chen drops the wet towel over his shoulder and walks toward Lin. No dialogue. Light: one warm-white pendant above the stove throws soft light down from upper right; cool rain light from the window fills the left edge.
[Shot 2] 5s: Close-up, Lin's upper body. Lin (voice: reference audio 1, low and slow) says: {I waited three years for this bowl.} After the last word her lips close; her knuckles tighten on the rim.
[Shot 3] 3s: Over-the-shoulder from behind Lin onto Chen. Chen stops half a step from the table, eyes dropping to the bowl. No dialogue.
[Shot 4] 3s: Close-up, Chen. Chen says: {It's still hot.} Lips close, shoulders ease a fraction.

Every line of dialogue ends with "lips close + one small action". Silent shots say "No dialogue" explicitly.

Dialogue duration table (English, ~2.5 words/s)

Words≤34–56–78–9>9
Shot length2.5 s3.5 s4.5 s6 ssplit the shot

Add 0.75 s per extra sentence-final punctuation mark. Too many seconds and the model repeats or stretches the line; too few and the tail of the sentence lands on the next person's face.

Eight self-contradictions to avoid

  1. "Takes two steps" and "feet planted" in the same line.
  2. "Speaks toward the back of the frame" while asking for a facial expression — the model turns them to camera.
  3. Lighting written as a landing spot ("light on her forehead") — you get a yellow blob.
  4. "Volumetric light / god rays" — Stable-Diffusion-era tokens; modern models want "a layer of smoke hangs in the air".
  5. "Hit, falls, spits blood" inside the swing shot — the result becomes the starting state.
  6. Large rigid-body motion (door kicked open, table flipped) finished inside one shot with a person — models render only start and end frames.
  7. Facial features for the back-to-camera person in an over-the-shoulder shot — a fourth person appears.
  8. Naming who "exits frame" — whoever you name, the model draws.

FAQ

How many cuts per segment?

3–6 shots in a 10–15 s segment is the stable range; beyond six most models start merging or skipping cuts.

Can a segment open on a silent establishing shot?

Yes, but keep it under a second: some models pull the voice track forward, and the longer the silence the larger the offset.

What about Chinese dialogue timing?

Calibrate at ~5 characters per second and use the same table with character counts.