Voice Cloning for Dubbing (and How to Do It With Consent)

Published October 11, 20267 min read

The biggest quality jump in dubbing in the last few years isn't more natural synthetic speech — it's the same synthetic speech attached to the right face. When a dubbed voice matches the person on screen, audiences stop hearing "a translation" and start hearing the video.

Quick answer

Speaker-preserving dubbing feeds a few seconds of the speaker's own line into a speech model, which re-speaks translated text in a matched voice. It works remarkably well — and it makes consent the central question of the craft. Clone only voices you have the right to use; for everyone else, use a designed voice, which takes minutes and clears nobody's likeness.

There is a moment in every multilingual test group when the dubbed version plays and nobody notices the switch. That's the whole effect: the same rhythm, the same register, a voice the audience already associates with the person speaking — in a language the person doesn't speak. This guide covers how that's produced, where it breaks, and how to do it without stepping on anyone's rights.

How a dubbing pipeline clones a voice

Modern dubbing doesn't train a voice model per speaker the way standalone cloning tools used to. It works per line:

  • The transcript carries word-level timestamps. This is the unglamorous hero of the whole process. Precise word timing lets the pipeline cut a clean reference window — the speaker's own line, from their own recording, a few seconds long.
  • The reference conditions generation. The model receives the reference audio and the new text, then produces speech in the target language shaped by that reference: pitch range, timbre, pacing habits.
  • Mode decides the performance. A steady mode keeps delivery consistent line to line — right for narration and interviews. An expressive mode follows each line's delivery more closely — right for acting, comedy, anything with dynamics.
  • Consistency comes from identity, not luck. The same speaker design is reused across every line they say, so line 47 doesn't drift into a different person than line 3.
SwiftDub voice settings: quality, timing and voice detail options plus the Stable, Dynamic and Custom voice modes, with one row per line and a Generate all button
The mode choice lives in one panel: Stable keeps a speaker's delivery consistent line to line, Dynamic follows each line's energy, and Custom swaps the clone for an original designed voice — no real person behind it.

The three ways cloned dubs fail

Having generated a lot of these, the failure catalog is short and specific:

1. Bad reference audio

If the reference window is quiet, noisy, or cut mid-word, everything downstream inherits the damage — the voice comes out unstable or thin. Pipeline-level fixes measure the reference before use and re-cut it; the user-level fix is simple: never accept a reference that sounds bad in the picker. If it sounds wrong there, every line will sound wrong.

2. The mid-sentence start

The subtle one. A reference that begins mid-sentence gives the model half a sentence to "finish" — in the original language. The result: the first fraction of a second of a dubbed line leaks the source language before switching. SwiftDub's quality check listens for exactly this on every clip and offers a fresh take, because the leak is invisible in waveforms and obvious to a human ear.

3. Lines that outrun the reference

A long, animated line delivered on a short, calm reference drifts: the voice stays similar but the energy doesn't match, and in dubbed dialogue that reads as "off" without anyone being able to say why. The practical fix is expressive mode for acted content, and splitting runaway lines instead of squeezing them.

Rule of thumb: if a single line sounds wrong, regenerate the line. If every line from one speaker sounds wrong, the reference is the problem — fix it once and everything downstream heals.

The consent question, taken seriously

Voice is identity, and the law is catching up fast. Tennessee's ELVIS Act (Ensuring Likeness Voice and Image Security), effective July 1, 2024, was the first US state law written explicitly for AI voice cloning: it extends the right of publicity to a person's voice and reaches the tools used to clone it. Other jurisdictions have followed with their own rules, and the direction of travel is consistent — a voice you don't have the right to use is a voice you don't clone.

The good news for video creators: the overwhelmingly common case is already cleared. If you're dubbing your own videos, or videos of people who agreed to be in them, the consent conversation is one sentence long — "I'd like your voice to work in other languages too" — and then it's documented. Where creators get into trouble is the middle cases: employees, interview guests, stock footage, family video.

A consent checklist that holds up:

  1. Written, specific, revocable. One paragraph naming the person, the use (voice cloning for translated versions of our videos), and the right to withdraw. Email counts. Nobody needs a lawyer for this.
  2. Scope it. "For dubbing our channel's videos into other languages" is scope. "For any purpose" is not something a reasonable person signs.
  3. Keep the record with the project. Consent and projects age at the same rate; store them together.
  4. Honor withdrawals. When someone opts out, stop using their clone and re-dub affected videos with a designed voice. Tools that keep per-line clips make this a mechanical swap rather than a re-run of everything.
  5. When in doubt, use a designed voice. An original synthetic voice has no likeness behind it — nothing to clear, nobody to notify. It's also the honest choice for documentary-style content about people who never signed anything.

Designed voices: the consent-free path

A designed voice is not a degraded clone. It's an original voice — configured once with age, pitch and character traits — used consistently for a speaker across every line. For interviews with strangers, corporate training content, or channels shipping forty videos a month across six languages, designed voices sidestep the entire consent question while keeping speaker consistency, which is most of what audiences actually need to follow who's talking.

In SwiftDub this is a first-class mode rather than an afterthought: pick the voice once per speaker, it stays identical on every line, and switching a speaker from clone to designed (or back) regenerates only their lines.

The bottom line

Voice-preserving dubbing is the difference between a translated video and the same video speaking another language. Use it wherever consent is clean — which, for most creators, is almost everywhere. Where consent is murky, a designed voice in three minutes is not a compromise; it's the professional answer. And whichever you choose, listen to the first second of every line — that's where both clones and pride go to leak.

Hear your own voice in another language

Import a clip and generate lines in each speaker's own voice — or design one. Free plan covers short clips end to end.

Open the studio

Frequently asked questions

How much source audio does a clone need?

A few clean seconds per speaker is the working minimum in modern pipelines — in dubbing, that sample is usually a single line from the video itself, cut at word boundaries. More helps, but the quality of the sample matters far more than the quantity.

Is voice cloning allowed on YouTube?

With consent, yes. Without it, you're exposed to right-of-publicity claims in a growing number of jurisdictions, plus platform rules against synthetic media of real people. Disclose realistic synthetic voices where platforms ask for it.

Can a cloned voice speak a language the person doesn't know?

Yes — that's the point. The model reproduces voice characteristics, not language ability. Pronunciation of language-specific sounds is generated, so an English speaker's clone can speak fluent Japanese.

What if the person has died — can their voice still be cloned?

Post-mortem rights vary by jurisdiction — some US states explicitly extend publicity rights after death. Absent a clear grant in a will or estate, treat it as unlicensed and use a designed voice instead.