4 ms·
I've been working in this area for a little over a year, so perhaps the artifacts in the generated speech are more apparent to me than to others. With that sai
by ansk 5y ago
I've been working in this area for a little over a year, so perhaps the artifacts in the generated speech are more apparent to me than to others. With that said, I don't find the displayed voices to be very compelling. All of them have artifacts which are apparent on a very localized scale - similar to compression artifacts, but I believe they are better described as phase discontinuities which result from the types of neural networks used to generate the waveforms in parallel. If you see a speech synthesis startup adding backing tracks to their speech samples, it's to cover up these artifacts. Removing these artifacts and improving the audio fidelity is a solvable problem, and you can look at the latest research for promising approaches. However, improving the naturalness and expressiveness of the speech at a high level is much further from being solved. I'm not caught up with the literature in this area entirely, but I have yet to see any examples of expressive text-to-speech that has natural intonation and prosody. The TTS researchers at Google have a bunch of research[1] which has essentially solved the task of monotone voice-assistant text-to-speech, but they haven't had as much success when it comes to more expressive speech. Best I've seen on this front is from an older paper[2], but the audio fidelity is quite poor there so it still leaves a lot to be desired.
[1] https://google.github.io/tacotron/ https://google.github.io/tacotron/
[2] https://audio-samples.github.io/#section-4 https://audio-samples.github.io/#section-4
- trompetenaccoun 5y agoAll of these sound really good to me, I think your ear being trained on these issues has a lot to do with it. We're very close to the point of it being indistinguishable from human speech, in fact with a few of these short clips I wouldn't be able to tell if I didn't know it's artificial speech.
- rasz 5y agoAll recordings are of low fidelity, like they sampled someone talking over GSM over voip over analog.
- chess_buster 5y agoI heard the artifacts too and was not impressed. I do not work with that industry.