Home › Glossary

AI Short-Drama Glossary: 20 terms explained

AI short drama, vertical micro-drama, novel-to-video, storyboard script, segments vs shots, character consistency, reference images, native audio, AI labeling and more — 60–120 words each.

AI short drama

A serialized show produced with AI generation rather than live shooting, typically vertical, 60–120 s per episode, dozens of episodes per title. Pipeline: parse the novel/script → character and scene reference art → storyboard → segment video → voice → compositing. Differs from 'AI video' by cross-episode consistency and series structure.

Vertical micro-drama

A phone-first 9:16 drama format: 1–2-minute episodes, a hook in the first 3 seconds, a cliffhanger per episode; monetized by per-episode unlocks or ads. China's market was ≈ ¥50.5B in 2024; internationally led by ReelShort and DramaBox.

Novel-to-video

The workflow that adapts a novel into a video series automatically. The real barrier is the parsing layer: episode splitting, appearance text an image model can reproduce, debris-free scene descriptions, a conflict and hook per episode. Tools without parsing are just video generators.

Storyboard script

Per-shot text a video model executes: framing and camera, who is in frame, action written as cause not result, dialogue with speaker, lighting as source direction and colour temperature, seconds. In AI drama the storyboard text is the prompt.

Segment vs shot

A segment is one video-model call, usually 10–15 s (30 s on Wan); a shot is a camera/framing change inside it, 3–6 per segment. Rule: long segments, short shots — fragmenting segments doubles cost, long shots drag.

Character consistency

Keeping one character's face, build and costume identical across shots and episodes. Achieved in four layers: reproducible appearance text, neutral-light white-background sheets, references injected per segment, and storyboards that list only who is in frame.

Reference image

An image passed with the prompt to lock a character's look, a scene's space or a prop's material. Limits vary (H3 ≤9, Qwen-Image ≤3, SeedDream ≤14); some models cannot combine a first frame with references.

Character sheet

A character's reference image: pure white background, frontal full body, key + fill + rim light, neutral expression, 9:16. The goal is recognisability, not beauty — any baked-in mood or lighting is inherited by every shot.

Scene plate

A people-free, debris-free scene reference: three layers of depth, materials, light source position. Test: 'if another scene were shot here, what shouldn't be there?' — broken mirrors, blood, crowds are plot debris and would appear in every episode before the event.

First / last frame

Images that pin a clip's opening or closing frame. First frames give continuity between segments, but on some models (e.g. H3) they exclude character references — continuity costs you the cast lock, the mechanical cause of blocking drift.

Native audio (audio-video generation)

The video model generates speech and ambience together with the picture, lip-synced to whoever speaks on screen. Supported by MiniMax H3, Seedance 2.x, Wan 3.0, Veo 3, Sora; others need post-production voice and lip-sync.

Reference audio

A 3–8-second voice sample bound to a character and sent with each segment so the voice stays consistent. H3 caps total reference audio at 15 s and 3 clips; matching from a preset library is two orders of magnitude cheaper than generating voices per character.

Dialogue timing

Converting a line into shot seconds by speech rate: ≈ 5 Chinese characters or 2.5 English words per second, +0.75 s per sentence-final mark. Too long and the model repeats or stretches; too short and the tail lands on the next shot.

Hook / cliffhanger

The per-episode ending suspense (identity reveal, unresolved emotion, reversal, rotated) and the attention grab inside the first 3 seconds. They decide completion and paid conversion, so the parser should produce them per episode.

Credit-based billing

How most AI video platforms bill: buy credits, deduct per call. Compare on cost per finished second or per series, and check expiry (Runway monthly credits don't roll over; Vidu subscription credits last 30 days).

720p / 1080p / 2K tiers

Video output tiers whose prices differ by 45–60%. Draft at the low tier, finalize at the high tier; vertical delivery is 1080×1920.

Prompt rewriting (prompt extend)

Automatic LLM expansion of your prompt (on by default on Wan, Qwen-Image and others). Turn it off for drama storyboards: the rewriter overrides cut points, durations and facing rules.

Content moderation

Compliance screening of prompts and outputs. Chinese models reject gore, corpses and nudity terms (e.g. Wan's DataInspectionFailed), usually without charge; keep violence at cause and consequence in storyboards.

AI-generated content label

Explicit or implicit labeling of AI-generated content. China's AIGC labeling measures require generators and platforms to label (mainland platforms watermark by default); international platforms require disclosure per their policies.

IAA / IAP drama

Two monetization models: IAP (pay to unlock episodes) and IAA (free with ads). Since 2025 IAA dominates in China, demanding higher output volume — the main entry point for AI production.