HomeLearn › AI Voice for Short Drama: Preset Voices, Voice Cloning and Native Audio — and Why Lip Sync Fails

AI Voice for Short Drama: Preset Voices, Voice Cloning and Native Audio — and Why Lip Sync Fails

There are three routes. Preset voice libraries match a character by gender, language and age at near-zero cost per character. Voice cloning is billed per voice slot or subscription and requires the speaker's consent plus an AI-content label. Native audio lets the video model speak the line itself: lip sync is inherent, but every regeneration is billed as video seconds. When lip sync fails, the model is rarely the cause. The script did not say who is speaking, the line was given the wrong number of seconds, or the reference voice's delivery overrode the performance note.

The three routes in one table

Short version: presets for supporting roles and narration, cloning for leads when the rights are in hand, native audio for shots with short lines. Most productions mix all three. Every pricing note below comes from the vendor's public page, checked 2026-09-02.

RouteExamplesBilling (official)StrengthCost of the choice
Preset libraryMiniMax system voices, CapCut Text to Speech, ElevenLabs premade voices, Azure standard voicesPer character: MiniMax speech HD $100 per 1M characters; ElevenLabs 1 character = 1 credit on v2 multilingual; CapCut page says freeInstant, auditionable, stable across episodesFinite catalogue; delivery style cannot be edited
Voice cloningElevenLabs IVC / PVC, Volcengine voice cloning, MiniMax rapid cloningPer voice (one-off or subscription) plus synthesis per character; ElevenLabs PVC from Creator at $22/month; MiniMax rapid cloning $3 per voiceUnique voice, reusable for the whole seriesConsent and labeling; PVC needs 30–180 minutes of audio
Native audioSeedance 2.0, MiniMax H3, Veo 3.1, WanNo separate line item: it is inside the per-second video priceLip sync is inherent; ambient sound comes freeDialogue regenerates with the picture; weak control over language and timbre

Route one: preset libraries are cheap, but you have to audition

Presets are billed per character, and drama dialogue is short. A 90-second vertical episode carries roughly 150–250 English words, about 1,000 characters. At the $100 per million characters MiniMax lists for HD speech on its pay-as-you-go page (source), an episode costs about ten cents. ElevenLabs bills credits: Creator is $22/month for 121,000 credits, and one text character equals one credit on the v2 multilingual models (pricing). CapCut's Text to Speech page lists 200+ voices filterable by gender, age, style and language and describes the tool as free (source). Azure standard voices cover 100+ languages and locales; note that each Chinese character is billed as two characters (Microsoft Learn).

Matching order is fixed: gender, then language, then age, then listen. In CapCut or the ElevenLabs library a person filters by tag and picks. Some end-to-end drama tools match for you: SceneMixer, for instance, assigns each parsed character a 3–8 second sample from MiniMax's system voice catalogue by gender, language and age, and lets the user re-match or upload. In SinCoSphere editorial testing, library matching costs roughly nothing per character; a bespoke generated voice costs a slot each.

Route two: cloning is a rights question before it is a price question

ElevenLabs documents two tiers. Instant Voice Cloning needs 1–2 minutes of clean audio and is available from the Starter plan ($6/month). Professional Voice Cloning needs 30–180 minutes, is available from Creator, requires identity verification, and may only be used on your own voice; fine-tuning takes 3–6 hours (docs). MiniMax lists rapid voice cloning at $3 per voice on its pay-as-you-go page. Volcengine's voice cloning is billed as prepaid voice slots, each allowing 15 training uploads with the newest overwriting the last (billing page); the annual fee is behind a login; a Volcengine developer-community article puts it near ¥150 per voice per year (community article, not an official price list). Azure custom voice is a limited-access feature you apply for.

Two legal facts for anyone shipping into or out of China. Article 1023 of the Civil Code extends portrait-right protection to a natural person's voice: no production, use or publication without consent (statute text). The AI-generated content labeling measures took effect on 2025-09-01 and cover synthesized audio (CAC). In practice: a written release from the voice actor that names the use, the term, and whether the clone may be reused on other titles.

Route three: native audio has the best lip sync and the most expensive retakes

The video model speaks the line, so the mouth and the sound come from the same pass. The switches differ by vendor. Seedance 2.0 takes a boolean generate_audio on task creation, supported on 2.0, 2.0 fast and 1.5 pro (Ark API). MiniMax H3 accepts reference audio, at most 3 clips and 15 seconds in total, and its first-frame image is mutually exclusive with reference images (docs). Google describes Veo 3.1 as a model for video with native audio; its prompting guide puts dialogue in quotation marks and describes sound effects and ambience as separate cues (Gemini API, guide). Alibaba's Wan guide (wan2.5–3.0) gives a formula: the speaker, the quoted line, then tone, pace, timbre and accent, with "No dialogue." as the explicit off switch (Model Studio).

Two costs. First, the line and the picture are one generation: a mispronounced word or a dragged syllable means regenerating the whole segment at per-second video rates. Second, language and timbre are controlled by prose. "A low middle-aged male voice" is a specific asset in a library and a probability distribution in a video model.

Emotion, gender, age: hard filters versus delivery

Gender, language and age are hard filters. Apply them first, then listen. The hard part is delivery: announcer, call-centre, cartoon. Delivery is a property of the voice itself, and no emotion note will remove it.

In SinCoSphere editorial testing, a male lead was matched to a news-anchor preset. Every line after that, including ones marked "lowered voice", "on the edge of tears" and "through clenched teeth", came out as a bulletin. A voice whose sample already had rise and fall made the same notes work at once. The delivery baked into a reference sample overrides the performance note. That is what to listen for when auditioning, more than resemblance.

Age must match the face. A character written as 25 with a settled 45-year-old sample makes viewers notice the mismatch before the story. Put emotion in the line note (tone, pace, pauses) and fix delivery by changing voices; do not try to do one with the other.

The three real causes of broken lip sync

One: speaker attribution. The model animates the mouth on the largest, nearest face in frame. A storyboard that says "Line: don't go" without naming who says it will, in any two-shot, move the wrong mouth. Fix: the speaker appears in the frame description first, the line names the speaker, and multi-person dialogue is cut into reverse shots, one face per shot.

Two: wrong number of seconds. SinCoSphere's calibration is about 2.5 words per second for English and 5 characters per second for Chinese. A ten-word line wants about four seconds. Give it two and the second half lands on the next shot's face; give it six and the model stretches syllables or reads the line twice to fill time. Drift is invisible at three seconds and obvious by twelve. See dialogue timing.

Three: rhythm in the reference clip. If the 15 seconds of H3 reference audio are read-aloud narration, the model copies the cadence and the character sounds like a script reading. Use conversational samples, not narration.

Routes one and two never had lip sync: TTS laid over a finished picture has no relationship to the mouth. Re-animate it with a lip-driving tool, or storyboard around it with reverse shots, over-the-shoulder angles and off-screen lines.

What voice actually costs per episode

  • Presets: usually bundled by end-to-end tools; called directly, about 1,000 characters of dialogue at $100 per million is roughly $0.10.
  • Cloning: a one-off slot or subscription per voice, plus synthesis per character; spread over a series it rounds to nothing.
  • Native audio: no line item, it is video seconds. Regenerating a 10-second segment once is paying for 10 more seconds. The real voice cost is your regeneration rate.

Voice is under 1% of a series budget; video generation is above 85% (see the cost breakdown). Do not save that 1% by skipping releases and labels.

A workflow you can copy

  1. At parsing, record each character's gender, age band and language in the character table.
  2. Match everyone from a preset library first; audition each 3–8 second sample specifically for delivery.
  3. Decide whether the leads justify a clone; get the release before the audio, and label the output.
  4. In the storyboard, give every line its seconds at 2.5 words per second, name the speaker, cut group dialogue into reverse shots.
  5. Use native audio for shots with lines of 8 seconds or less; use TTS plus reverse shots or off-screen delivery for long speeches.
  6. Before export, listen through the whole episode for consistent delivery per character and age that matches the face.

For how lines and seconds are written into a shot, see video prompt writing; terms are defined under native audio and reference audio.

Key figures and sources

FigureValueSourceDate
MiniMax speech HD, pay-as-you-go$100 per 1M characters; rapid voice cloning $3 per voiceplatform.minimax.io2026-09-02
ElevenLabs Creator plan$22/month, 121,000 credits; Professional Voice Cloning from this tierelevenlabs.io2026-09-02
ElevenLabs cloning audio requirementsInstant 1–2 minutes; Professional 30–180 minutes plus identity verificationelevenlabs.io2026-09-02
MiniMax H3 reference audio limitAt most 3 clips, 15 seconds totalplatform.minimaxi.com2026-09-02
China AI-content labeling measures in force2025-09-01www.cac.gov.cn2026-09-02

FAQ

Will preset voices sound alike across characters?

Yes, especially for same-gender, same-age roles. Give leads voices that differ in pitch, pace and breathiness; supporting roles can share.

How much audio does a clone need?

ElevenLabs: 1–2 minutes for Instant, 30–180 minutes plus identity verification for Professional. Volcengine sells prepaid slots with 15 training uploads each. Clean audio only: no music, no room reverb.

Can I clone a celebrity or influencer voice?

No. China's Civil Code Article 1023 protects voice like portrait rights; ElevenLabs' professional tier accepts only your own verified voice. A hired actor still needs a written release, and the audio needs an AI-content label.

Which is cheaper, native audio or TTS?

Per character, TTS by a wide margin. Native audio has no separate price: it is the per-second video rate, paid again on every regeneration. It pays off on short lines with a high first-pass success rate.

Why is the mouth always half a beat late?

Usually too many seconds. Re-time the line at 2.5 words per second and hand the spare seconds to action or a reaction shot instead of letting the model fill them.

One character, two languages: same voice?

Pick a target-language voice of the same gender and age band. Do not let the source-language voice read the translation; the accent will be wrong. Multilingual clones need every line checked for accent drift.

Sources

Tools mentioned

Read next