4 ms·
Impressively done! It sounds like you're doing 1) doing voice recognition with voice time clues, which Whisper and the like provide, breaking it up into sente
by waldrews 3y ago
Impressively done! It sounds like you're doing
1) doing voice recognition with voice time clues, which Whisper and the like provide, breaking it up into sentence (or similar) units; you don't need to time match individual words, but you need to time match at coarser grain.
2) using a translation engine that allows for multiple alternative translations
3) cloning the original voice, regardless of language
4) choosing the translation that has the best time match (possibly by syllable counting, or by actually rendering and timing the translations). If there isn't a close translation, maybe you're asking ChatGPT to forcibly rephrase?
5) Maybe some modest pitch-corrected rate control to pick out path that gets you closest to the timing?
Did I get any of that right?
- euazOn 3y agoI also noticed that the third sample with Chinese sounds slightly sped up in the first English segment, so there may be also an element of postprocessing the dub (speeding it up/slowing it down).
- leobg 3y agoYes. Though I don't like this solution. It breaks the flow. And it also doesn't really fully solve the problem. Overruns still accumulate if they happen too frequently. One second here, one second there... the further you get into the video, the worse it gets. I think it would be better to either slow down the underlying video or solve the overrun issue on the translation level. A good professional dubber will find translations that will even out in terms of timing. That's something an AI should be able to do better instead of worse.
- odiroot 3y agoThe last sample from BBC is really hilarious when translated to Polish. Something definitely went wrong and the voice speaks like a drunkard.
- waldrews 3y agoOoh and you're probably doing a split into voice and non-voice tracks of the original, and keeping non-voice at original volume, but lowering the voice track.
- davidzweig 3y agoI think it's a speech to speech model, I know about seamlessm4t: https://www.google.com/amp/s/about.fb.com/news/2023/08/seamlessm4t-ai-translation-model/amp/ https://www.google.com/amp/s/about.fb.com/news/2023/08/seaml... Interesting, but what inference engine supports it to run at decent speeds?
- leobg 3y agoVery good! Yes, that's basically how it works. I don't do any pitch-correction. But I do check the TTS output for lenght, and I re-generate if it doesn't match my time contraints. I also have an arranger that tries to figure out when to play an utterance early (i.e. earlier than in the original) in order to make up for the translated version being longer. I try to make the translations match the speaker's character, as well as the context. So ideally, Alex (Sample 2) will still say "Salut" even in German (instead of translating that greeting, too). And I need to monitor for speaker changes. This is because I can't clone the voice unless I have a decent amount of sample data. If Elon just says "Yes", cloning the voice based on just that one syllable will make it sound like a robot. But I also can't just blindly grab any voice around it, since that might be somebody else's voice.