This is the operation that makes dubbing possible. It has a name in the industry, it relies on a specific technology, and it has limits worth knowing before you get frustrated with an imperfect result.
In a professional dubbing studio, nobody removes a voice from a film. They ask the producer for a version of the soundtrack without dialogue, called M&E: music, effects, ambiences, everything but voices. Actors record over it, and the mix reassembles a complete film in the new language.
That is why a dubbed film keeps the exact same explosions, the same footsteps in gravel and the same score as the original. Nothing was recreated. Only the voice layer was replaced.
For a video you found online, that track does not exist. It has to be rebuilt from the final mix. That is what source separation does.
A separation model was trained on tens of thousands of tracks where both the mix and the isolated stems were available. Through enough examples, it learned what a human voice looks like in a signal, and what does not.
Given an unknown clip, it slices the sound into very fine frequency and time cells, decides for each one whether it is voice or not, then rebuilds two files from those decisions.
The key point fits in one sentence: when a voice and a piece of music occupy exactly the same frequency at the same instant, the information is lost. It is not hidden. It no longer exists in the mix. No model, however good, can recover it. It can only guess.
Those guesses produce the defects you hear:
What hurts the result most is not the model, it is the source. Short form social media clips go through aggressive compression and loudness maximised mixing where everything overlaps. On that material even the best models leave artefacts. A dialogue scene with restrained music and a clean source comes out far better.
| Type of clip | Separation quality |
|---|---|
| Calm dialogue, two characters | Excellent |
| Tense confrontation, moderate music | Good |
| Battle scene, orchestral score at full volume | Poor |
| Music video or sung passage | Very poor |
When you submit a video, Dubba builds the music and effects track automatically and keeps it. That is what you hear behind your voice in the final dub. In parallel, the isolated voice track is used to transcribe the dialogue and identify the different characters, which feeds the rythmo band.
Try it on a videoA film's full soundtrack minus the dialogue. The standard delivery format between a producer and a dubbing studio.
Almost, never entirely. Where voice and music overlap exactly, the information is gone.
Those are the model's hesitation zones, concentrated in the high frequencies and most noticeable in quiet passages.
Only partly. On a heavily compressed source, the difference between a decent model and a state of the art one is often inaudible. The limit comes from the material, not the tool.
Technically separation produces both. The voice track is what powers transcription and character detection.