HomeLearn › Subtitles, Dubbing and Localization for Micro-Drama Going Global: Text Expansion, Lip-Sync Drift and Cultural Adaptation

Subtitles, Dubbing and Localization for Micro-Drama Going Global: Text Expansion, Lip-Sync Drift and Cultural Adaptation

Localization is three layers, not one. The script layer changes names, forms of address, settings and premises, and must be done at parsing so the changes flow into character sheets and storyboards. The subtitle layer can only change words; it cannot fix the Chinese shop sign in frame. The dub layer meets text expansion: Chinese-to-English dialogue takes 30–50% longer to say, so mouths and seconds drift apart. Handle it in order: compress the translation, split the cue, re-time the shot, and only then regenerate.

Start here: three layers, three different fixes

Treating "translate the subtitles" as localization is the most common source of rework in exported micro-drama. Cultural information sits on three layers with different timing and cost.

LayerWhat changesWhenCost orderTypical trap
ScriptNames, address terms, locations, festivals and currency, institutional premises (bride price, household registration)At parsing, so it flows into character, scene and shot tablesOne LLM call, a few dollars at mostChanging it after parsing means redoing every reference image and storyboard
SubtitleWording: translation, units, equivalents for untranslatable address terms, cue position and line countAfter picture lockMachine translation near zero; human review per episodeCannot touch Chinese signage, couplets or banknotes in frame
DubTarget-language voice, line seconds, lip syncDub over the finished picture, or re-time and regenerate dialogue shotsAI dubbing per minute; regeneration per video secondText expansion drifts the mouth; a source-language voice has an accent in the target language

One question decides which layer owns a problem: is it in the picture? Anything visible (signs, money, costume, buildings) is script-layer only; anything that lives only in dialogue is a subtitle job.

Subtitle specs: characters per line, reading speed, the 9:16 safe zone

TikTok and YouTube Shorts publish no subtitle reading-speed rule. The industry reference is Netflix's timed-text style guides, per language (English, Simplified Chinese, Japanese).

LanguagePer lineMax linesReading speed (adult / children)
English42 characters220 cps / 17 cps
Simplified Chinese16 characters29 cps / 7 cps
Japanese (horizontal)13 full-width characters24 cps

Those are 16:9 ceilings. On a 1080 px wide vertical frame at a 56–64 px font, a line holds about 30 English or 12–14 Chinese characters: plan for 70–80% of the Netflix limit.

Safe zones likewise have no single official page; third-party template guides converge on these 1080×1920 numbers: TikTok's UI covers roughly the top 108–150 px, bottom 320–440 px and right 120–180 px; YouTube Shorts about the top 120 px and bottom 300 px (guide one, guide two, neither from a platform). Working rule: burn subtitles at 65–75% of frame height, centred, two lines maximum; keep the bottom quarter and the right 180 px free of text.

Text expansion: why the English version's mouths are wrong

Chinese to English keeps the meaning and lengthens the speech. SinCoSphere's calibration: about 5 Chinese characters and 2.5 English words per second. 「你根本不知道我为你放弃了什么。」 is 15 characters, three seconds. Literal: "You have absolutely no idea what I have given up for you," 12 words, 4.8 seconds, 60% longer. Compressed: "You don't know what I gave up for you," 9 words, 3.6 seconds, 20% longer. The practical band is 30–50%.

Where the expansion lands depends on the dub. Overlaid: the English audio runs a second past the Chinese mouth, so either speed it up (above about 1.15x it stops sounding human) or let it spill onto the next shot. Native audio: regenerate the dialogue shot in English; the mouth is right and the shot is billed again by the second.

Three fixes, in order. Compress: cut words, not meaning; drop fillers like "absolutely", break subordinate clauses into short sentences. Split: two cues on one shot, half the seconds each, which brings reading speed back under 20 cps. Re-time: give the line 1–1.5 seconds more in the storyboard and regenerate that shot, rather than pushing subtitles to 25 cps. When several languages ship together, time the storyboard to the longest one and let the others hold on a reaction. See dialogue timing.

AI dubbing: three options and what vendors say they cost

End-to-end localization services. A Tencent Cloud Media AI article claims 100 episodes into 9 languages (900 finished episodes) for about ¥2,433, roughly ¥2.7 each, in 48 hours. Its per-minute items: watermark removal ¥3, ASR ¥0.03, LLM translation ¥0.60, OCR ¥0.80, AI dubbing ¥0.5–9 by mode, subtitle rendering ¥0.063, compliance review ¥0.08; it puts human translation at "over ¥200 per episode per language" (source; vendor claim, not verified by SinCoSphere).

Dubbing platforms. ElevenLabs bills Dubbing in credits per source minute: automatic with watermark 2,000, without 3,000; Dubbing Studio 5,000 and 10,000. On Creator ($22 for 121,000 credits) that is about 40 minutes of unwatermarked automatic dubbing, roughly $0.55 per minute (source).

Inside the editor. CapCut lists AI Dubbing and a video translator, with auto-captions in 20+ languages (source); the dubbing language list changes. Fine for a pilot, not for a 100-episode batch.

None of the three solves lip sync; any overlaid dub needs the fixes above. For choosing voices and clearing clone rights, see AI voice for short drama.

Cultural adaptation: script layer or subtitle layer

Sort every adaptation by whether subtitles can carry it.

  • Subtitle layer can handle: names (transliterated or replaced), units and currency, equivalents for address terms with no counterpart ("CEO" for 总裁, a first name for 师兄), idioms swapped for target-language idioms.
  • Only the script layer can handle: spaces (courtyard house, hutong, dorm); visible props (spring couplets, red envelopes, banknotes, payment QR codes); institutional premises (bride price, household registration, a state job); plot devices like the gaokao that need lived experience.

Script-layer changes belong at parsing, because character looks, scene descriptions and prop lists are parsing outputs that then flow into reference images and storyboards. Some tools expose this as a setting: SceneMixer offers target-region presets at parsing (US/EU, East Asia, Southeast Asia, Latin America) that rewrite names, ethnicity in appearance descriptions, street and building types and visible on-screen text, with UI language and output language set independently. Whether LibTV or Xiaoyunque offer an equivalent is not stated on their sites. Without such a setting, prepend a "target market" note before parsing, or edit the character and scene tables by hand and regenerate the reference images.

One case: the male lead is a "group chairman". In subtitles, "CEO" is fine. But if the plot runs on inheriting a family conglomerate, a US audience accepts it, a Southeast Asian audience recognizes it, and the Latin American equivalent may be a different family structure. That is a script-layer premise decision, not a translation problem.

Lip-sync drift after dubbing: three causes, plus one for export

The three general causes — speaker not named, wrong seconds, reference-voice delivery overriding the performance note — are covered in the voice article. Export adds a fourth: target languages have different speech rates. A three-second shot holds 15 Chinese characters, 7–8 English words, fewer still in syllable-heavy Spanish. Language versions cannot share one set of storyboard timings.

The routine: the primary language (usually English) sets storyboard timing; other languages are compressed at the subtitle layer; close-up dialogue shots are regenerated once per language with native audio; wide shots, reverse angles and off-screen lines take an overlaid dub. That leaves 2–4 regenerated shots per episode.

QA checklist before export

  • Every cue two lines or fewer; 36 English or 14 Chinese characters per line (80% of the Netflix ceiling).
  • Reading speed: English 20 cps, Chinese 9 cps; split cues that exceed it.
  • Subtitles at 65–75% of frame height; no text in the bottom quarter or the right 180 px.
  • One name table per series; every name consistent across episodes.
  • Address terms consistent: one rendering per relationship.
  • Visible Chinese text (signs, phone screens, documents) handled at the script layer; otherwise regenerate or mask the shot.
  • Dub speed 1.15x or less; faster goes back to compression or re-timing.
  • Check every close-up speaking shot for lip sync; failures go on the regeneration list.
  • Run each episode past the target market's review lines (religion, politics, minors); a China release still needs the AIGC label (labeling measures).
  • Show three episodes to a native speaker cold; ask which line they missed and where they were pulled out.

Market data: which market first

Sensor Tower's Q1 2026 report: over 850 million downloads, up 140% year on year; in-app purchase revenue about $750 million, up 20%; Southeast Asia at 32% of global downloads, Latin America 23%, India 22%; Southeast Asian users average about 40 minutes a day against 25 globally (source). DataEye's estimate via 36Kr, counting purchases, subscriptions, ads and off-app content: the overseas short-drama market reached about $4 billion in 2025, with over $5 billion expected in 2026 (source).

Revenue leaders remain DramaBox and ReelShort at roughly $140 million each for the quarter, with paying users concentrated in English-language markets; download growth is in Southeast Asia, Latin America and India. So: English first for revenue; Indonesian, Thai, Vietnamese and Filipino for download growth; Spanish and Portuguese for Latin America. Channels are in distribution channels; the format playbook is ReelShort-style vertical drama with AI.

Key figures and sources

FigureValueSourceDate
Southeast Asia share of short-drama app downloads, Q1 202632% (Latin America 23%, India 22%)sensortower.com2026-09-02
Global short-drama app downloads, Q1 2026850M+, up 140% YoY; IAP revenue about $750Msensortower.com2026-09-02
Overseas short-drama market size, 2025 (DataEye)About $4 billion36kr.com2026-09-02
AI localization cost, Tencent Cloud Media AI (vendor claim)100 episodes x 9 languages about ¥2,433; about ¥2.7 per finished episodecloud.tencent.com2026-09-02
ElevenLabs Dubbing credit costAutomatic 2,000–3,000 credits per minute; Dubbing Studio 5,000–10,000elevenlabs.io2026-09-02
Netflix English subtitle spec42 characters per line, 2 lines, 20 cps adultpartnerhelp.netflixstudios.com2026-09-02

FAQ

How many characters per subtitle line on a vertical drama?

About 30–36 English or 12–14 Chinese characters, two lines maximum, at 65–75% of frame height. Split anything longer into two cues.

How much longer is English than the Chinese original?

Spoken, usually 30–50% longer. At 5 Chinese characters and 2.5 English words per second, a 15-character line rendered literally as 12 words needs 1.8 extra seconds. Compress, then split, then re-time.

What does AI dubbing 100 episodes cost?

Tencent Cloud Media AI claims about ¥2,433 for 100 episodes into 9 languages in 48 hours (vendor figure, unverified). ElevenLabs: 3,000 credits per source minute for unwatermarked automatic dubbing. Lip sync is extra either way.

Should Chinese character names be changed?

Western audiences retain transliterated Chinese names poorly, so renaming is standard; Southeast Asian markets with large Chinese-heritage audiences can keep them. One name table per series.

What about Chinese shop signs in the picture?

Subtitles cannot fix them. Set the target market at the script layer and regenerate the scenes, or mask and replace in post. Series built for export should pick the target region at parsing.

Can multiple language versions share storyboard timings?

No. Time the storyboard to the longest language (usually English or Spanish), compress the others in subtitles, regenerate close-up dialogue shots per language, and overlay a dub on wide and off-screen shots.

Sources

Tools mentioned

Read next