How Small Studios Run an AI Short-Drama Pipeline: Roles, Versioning, QC Gates and 3–5 Titles a Month
Three to five titles a month at 60–80 episodes each means shipping 8–18 episodes every working day. Throughput is set by two numbers, not by the model: reviewer hours (30–60 minutes per episode) and per-account concurrency (3–10 minutes per segment, 50–70% usable first pass). One 5-task seat yields segments for 15–20 episodes a day; one reviewer clears 8–16 — so a 4–6 person studio needs two reviewers. Versioning and QC gates keep rework from eating that capacity (typical, SinCoSphere editorial testing).
Throughput math: how many episodes a day
Arithmetic before the org chart. A 90-second episode is 6 segments of 15 s. A segment takes 3–10 minutes — call it 6 — and 50–70% are usable first time, so each usable segment costs 1.4–2 calls; call it 1.7, about 10 calls per episode. A seat with 5 concurrent tasks clears 10 calls in two waves, roughly 12 minutes; with queueing and hand-offs budget 20–25 minutes per episode — 15–20 episodes in an 8-hour day.
Review runs 30–60 minutes per episode, so one reviewer clears 8–16 a day; storyboard review at 20–30 minutes, 16–24 a day. Three to five titles at 60–80 episodes is 180–400 episodes a month, 8–18 per working day. Working backwards: one generation seat, one or two reviewers, one storyboard editor, half an artist, one editor-distributor. Capacity is a reviewer number, not a model number — the sum most studios get backwards (typical, from SinCoSphere editorial testing).
Six roles, six QC gates
Each role owns one gate; pass criteria are checkboxes, never "looks good".
| Role | Input | Output | QC gate | Pass criteria |
|---|---|---|---|---|
| Story lead / parse reviewer | Licensed IP or original script | Approved cast, scene and prop sheets, episode outline | G1 Parse | Appearance reproducible (face, brows and eyes, hair, build, one anchor); conflict and hook in every episode; ≤4 locations; episode count and length fit the channel |
| Art lead | Approved cast and scene sheets | Character sheets, scene plates, prop images, look table | G2 Art | White background, frontal full body, neutral light; features match the sheet; plates empty with an explicit light source; every look tagged with its first episode |
| Storyboard editor | Outline + asset library | Per-episode storyboard with a version number | G3 Storyboard | Segments ≤15 s with 3–6 shots; dialogue timed at ~2.5 words/s; only in-frame people named; no banned terms, no contradictions |
| Generation and picking | Locked storyboard version | One active take per segment | G4 Footage | Face, wardrobe and set continuous with the previous segment; lip-sync on the speaker; cuts executed; no stray extras or watermarks; first-pass usable rate ≥50% |
| Editor | Active takes + voice + subtitles | Finished episode | G5 Cut | No jumps between segments; levels consistent; subtitles verbatim; 60–120 s; hook inside the first 3 s |
| Distribution | Episode + cover and title | Published episode + metrics | G6 Release | AI disclosure set; rights file on record; authorship log filed; title and cover clean; schedule updated |
Smaller teams merge with care: story with storyboard, art with editing; generation and picking must never double as review — people pass their own takes.
Versioning: storyboards and segments
Rework eats capacity when versions are unclear: a storyboard changes and nobody knows which segments to redo; a segment is redone and nobody knows which take is in the cut. Three rules.
- Storyboards carry a version. v1 is the AI draft, human text edits v1.1 and v1.2, an AI regeneration v2. Lock the version before generating and stamp every segment "from storyboard vX".
- Segments keep history, one active take. Compositing reads the active take only; switching the active take is free and involves no regeneration.
- A storyboard change touches only the affected segments. A line change in segment 3 redoes segment 3; a character look change redoes every segment from the look's first episode — which is why looks are frozen at the art gate.
What is free and what bills again (common platform terms; check yours): storyboard text edits free; an AI redraft of an episode's storyboard billed again per episode; segment regeneration full price; switching takes free; compositing billed by duration; voice — native-audio models take the reference clip at generation time, so a voice change is a regeneration, while post-dubbing changes it for free. SceneMixer-style end-to-end platforms keep storyboard versions and segment takes in one project, each take stamped to its storyboard version; a single-purpose stack (Kling, Hailuo or Runway plus CapCut) keeps that mapping by hand in a sheet, and one missed row is an episode nobody can reassemble.
QC gates: the return rules
A gate is defined by what it sends back, not by its checklist.
- G1 back to story: adjectives instead of features ("handsome", "cold"), no hook at the episode end, more than 4 locations.
- G2 back to art: mood lighting or a non-white background, features off the sheet, people or plot debris in a plate, a look without a first-episode tag.
- G3 back to storyboard: a segment over 15 s or 6 shots, dialogue over the timing table, off-screen characters named, lighting as a landing spot, banned terms.
- G4 back to generation: continuity break, lip-sync on the wrong person, cuts ignored, stray extras, watermark. Under 50% usable for two episodes running, the episode goes back to G3, not to a rerun — the problem is in the text.
- G5 back to the editor: jumps between segments, uneven levels, subtitle typos, length outside 60–120 s, hook later than 3 s.
- G6 back to distribution: AI disclosure, rights file or authorship log missing; title or cover breaking channel rules.
Record every gate's pass rate. G4's first-pass rate is the pipeline's key number: when it drops, look at G3 first.
Cost control: three principles and one sum to get right
Three principles, ordered by how much they save.
- Review the storyboard before generating. Text edits are free; regeneration is full price. Ten more minutes at G3 saves a round at G4 — about 30–40% of an episode's generation spend.
- Push segments toward 15 s. The same 90 seconds in 6 segments is half the calls of 12, and the cuts happen inside the segment.
- 720p drafts, 1080p finals — but do the sum. The low tier is usually 45–60% of the high tier's price. If delivery must be 1080p, drafts are an extra cost: at half price and a 60% first-pass rate, straight-to-1080p costs about 1.7 calls per usable segment; 720p draft then 1080p final costs about 0.5 × 1.5 + 1.2 = 1.95 call-equivalents — dearer. Drafts pay off in three cases only: the channel accepts 720p (on a phone the tiers are hard to tell apart), high-risk segments under a 40% first-pass rate (action, crowded frames), and a pilot's first three episodes. Never render a whole title twice.
Order of magnitude: at about $60–110 per 10 episodes in 720p and $130–230 in 1080p, a 60-episode title is roughly $360–660 or $780–1,380 in generation credits; 3–5 titles a month land between $1,100 and $7,000 depending on tier and regeneration rate (typical, SinCoSphere editorial testing). invideo's own case claims $1,000 per episode with three people over three days (vendor claim, July 2026) — a labor-inclusive number, not a credit number.
Tooling by stage, not by brand
Both routes scale; they differ in who maintains the version mapping.
| Stage | End-to-end route | Single-purpose stack | What matters |
|---|---|---|---|
| Project management | Project / episode / segment status inside the platform | Notion, Airtable, Trello | A status column for all six gates |
| Asset library | Cast / scene / prop library in the project | Google Drive with a naming rule; Frame.io for review | Look table carries first-episode tags |
| Storyboard | Generated in the platform, versioned | LLM + segment template + a version column in the sheet | Lock the version before generating |
| Video generation | LibTV, LTX Studio, invideo AI | Kling, Hailuo, Runway, Pika, Veo via Flow, Vidu | One model per episode, never mixed |
| Editing | One-click compositing | CapCut, Premiere, DaVinci Resolve | Cloud concatenation beats local re-encoding |
| Delivery | Export 1080×1920 vertical | ReelShort, DramaBox, YouTube Shorts, TikTok | Each channel's disclosure flow differs |
Seven common organizational failures
- Reviewing your own takes. G4 pass rates inflate and problems surface at G5, where rework costs double.
- Storyboarding before art is locked. One revised character sheet redoes every segment from that look's first episode.
- Storyboard changes by word of mouth. No version number, so after a regeneration nobody knows which take is in the cut.
- Several titles sharing one account's concurrency. The urgent title cannot get a queue slot; assign seats or accounts per title.
- Managing the regeneration rate as a KPI. "Good enough" reaches the upload and the audience does your QC in the comments.
- A broken rights chain. The license omits AI derivative works or the app's territories, discovered the day before release.
- No naming discipline in the library. The wrong look reaches episode 20 and the lead is back in episode 3's coat.
Rights, disclosure and the authorship log
Copyright. The US Copyright Office's Part 2 report (29 January 2025) holds that prompts alone do not make AI output copyrightable; the human-written script, human edits and the arrangement of segments into an episode can be registered, with AI-generated material disclosed. For a studio that means an authorship log per episode: who wrote the script, which storyboard version was human-edited, which takes were picked and cut. The pipeline's version records are that log.
Licensing and talent. The IP license must cover AI adaptation, derivative works and every territory the app sells in — ReelShort-style apps sell worldwide. No real person's face or voice as a reference; right-of-publicity claims are state law. Set the platform's AI-content disclosure on every upload.
Key figures and sources
| Figure | Value | Source | Date |
|---|---|---|---|
| invideo first-party case (vendor claim, 2026-07-24) | $1,000 per episode; 10 episodes; 3 people; 3 days | invideo.io | 2026-09-02 |
| MiniMax Hailuo H3 list price | $0.08/s at 768p, $0.13/s at 2K | platform.minimax.io | 2026-09-02 |
| Wan 3.0 list price on Alibaba Model Studio | ¥0.6/s at 720p, ¥1.2/s at 1080p | help.aliyun.com | 2026-09-02 |
| Seedance 2.0 per-second price (press report, 2026-03-04) | about ¥1 per second | finance.sina.com.cn | 2026-09-02 |
| US Copyright Office, AI report Part 2 (29 Jan 2025) | Prompts alone do not make AI output copyrightable; human-authored elements can be registered | www.copyright.gov | 2026-09-02 |
FAQ
Can three people ship three titles a month?
Yes, at 60 episodes each in 720p: one on story and storyboard review, one on art plus editing and release, one on generation and picking; the first two cross-review footage. Past three titles, hire a reviewer first.
How many generation seats do we need?
Episodes per day divided by 15. Even five 80-episode titles a month need only two 5-task seats; spare seats are useless if review cannot keep up.
Does one changed line mean regenerating the whole episode?
No — only that segment, provided every take is stamped with the storyboard version it came from.
What first-pass usable rate should we target?
60% is normal. Under 50% for two episodes running, send the episode back to the storyboard gate; over 80%, storyboard review can relax.
Can we mix two models inside one episode?
Not advised: mismatched encoding forces a re-encode at compositing and the texture jumps. Fix one model and one tier per title.
Should we build our own node-based workflow?
Not below three titles a month. A home-built pipeline's real cost is maintaining the version mapping and the asset library — exactly where volume production breaks.
Sources
- https://invideo.io/blog/how-to-make-ai-micro-drama/
- https://platform.minimax.io/docs/guides/pricing-paygo
- https://help.aliyun.com/zh/model-studio/model-pricing
- https://finance.sina.com.cn/tech/roll/2026-03-04/doc-inhpvpqk3838479.shtml
- https://www.copyright.gov/ai/