3 ms·
My guess would be that text-to-speech scales very well for arbitrary data, for e.g. automatic audiobook generation, speech-to-speech does not. But yeah, fully
by schroeding 4y ago
My guess would be that text-to-speech scales very well for arbitrary data, for e.g. automatic audiobook generation, speech-to-speech does not.
But yeah, fully agreed, for individual projects speech-to-speech appears to be a better idea, much more data to work with in there. Otherwise it will be a Vocaloid-like experience, where you have to tinker with the intonation of individual words.
There is significant work in this area, too, e.g. Zero-Shot Voice Style Transfer: https://auspicious3000.github.io/autovc-demo/ https://auspicious3000.github.io/autovc-demo/