HomeGuides › Character Consistency in AI Video: four layers that stop faces drifting across episodes

Character Consistency in AI Video: four layers that stop faces drifting across episodes

Face drift is rarely 'the model is bad'; it is usually a broken reference chain. Stable pipelines use four layers: ① parse appearance into reproducible specifics (face shape, brows, eyes, nose, lips, skin tone, hair, build, apparent age plus one anchor — never 'handsome'); ② shoot the character sheet on white, frontal, neutral light, shoes seen from the side (kills high-angle framing); ③ pass the full-body sheet and a face crop as reference images with every segment; ④ in the storyboard, list only who is in frame and never name who is out. Drop a layer and consistency drops a grade.

Layer 1 — Appearance an image model can read back

"A cold, striking man around thirty" cannot be reproduced. "Long narrow face, straight sharp brows, hooded eyes, high nose bridge, thin lips, cool pale skin, short black hair combed back, 1.85 m and lean, a 1 cm old scar at the tail of the left brow" can. Memory anchors must be static (scar, mole, jewellery, hairstyle), never an expression — write "eyes crinkle when he smiles" and the sheet grins, then every segment opens grinning.

Do not describe stance or gait either: "leans forward" in a full-body shot means head-near-feet-far, i.e. a high-angle render.

Layer 2 — Six hard rules for the character sheet

  • Pure white background (its own clause plus a negative list; buried in another paragraph it gets overridden)
  • Frontal full body, level lens, camera at half the subject's height
  • Directional key + fill + rim light; no flat light (flat light = plastic skin)
  • Priority face > posture > costume; never widen the shot to show clothes
  • Neutral expression, no 'cinematic' grade
  • 9:16 at 2K-class resolution so a face crop can serve as a second reference

Layer 3 — Reference injection

Send at least two images per character per segment: the full-body sheet and a face crop; one per person in multi-character frames, plus the scene plate. Mind model caps: MiniMax H3 accepts ≤9 references and surcharges beyond five; Wan's reference media are free. On some models first-frame and reference images are mutually exclusive — continuity costs you the cast references, which is the mechanical cause of cross-segment blocking drift; mitigate in the storyboard with scene anchors.

Layer 4 — Continuity discipline in the storyboard

  • Only list who is in frame; never write "others exit" — naming raises intrusion probability.
  • Actions completed in the previous segment become a stated end-state, not a replay.
  • Temporary visible states (wounds, costume changes, altered props) are restated every segment; set-dressing lines carry no bare nouns.
  • Positions use scene anchors ("the last bench by the window") not frame directions ("screen left") — a camera change invalidates the latter.

FAQ

Same reference, very different fidelity across models — why?

Models weight reference images differently and understand culture-specific costume terms unevenly. For period Chinese costume, Chinese-native image models produce more faithful sheets.

Can face-swap fix it in post?

Partially, but it costs more than getting the sheet right, and fails on profiles, occlusion and crowded frames.

How do I handle a character with several looks?

Keep multiple sheets (looks) for the same character bound per episode; keep one shared facial description and only vary costume and hair.