4 ms·
I don't understand why text to speech approaches are so common. It's really hard to specify exactly what you want with text. It seems to me like speech-to-spe
by welshwelsh 4y ago
I don't understand why text to speech approaches are so common. It's really hard to specify exactly what you want with text.
It seems to me like speech-to-speech would be much better: start with your best attempt to produce the audio yourself, with the emotion, rhythm and timing you want. Then let the AI do the "last mile" transformation, taking your voice and making it sound like someone else, like how neural style transfer can change a picture to another style.
- schroeding 4y agoMy guess would be that text-to-speech scales very well for arbitrary data, for e.g. automatic audiobook generation, speech-to-speech does not. But yeah, fully agreed, for individual projects speech-to-speech appears to be a better idea, much more data to work with in there. Otherwise it will be a Vocaloid-like experience, where you have to tinker with the intonation of individual words. There is significant work in this area, too, e.g. Zero-Shot Voice Style Transfer: https://auspicious3000.github.io/autovc-demo/ https://auspicious3000.github.io/autovc-demo/
- staindk 4y agoI'm late to this but IIRC this is kind of the tech that LTT has started making use of for spanish audio - I don't know any of the nitty gritty details but I think they feed the english track as well as the script into the AI and get a much more natural-sounding translation out of it. For sections that don't come out "right" you can help it along by re-training just that section etc. See vid for some discussion around it - https://www.youtube.com/watch?v=_5uCvcyD0Eo https://www.youtube.com/watch?v=_5uCvcyD0Eo