3 ms·
Showing one of the top two samples from their blog (full prediction) along with the one you link (only acoustic model) would more clearly show what you are gett
by kastnerkyle 10y ago
Showing one of the top two samples from their blog (full prediction) along with the one you link (only acoustic model) would more clearly show what you are getting at in the explanation, since how the text inference works is most of the complexity in this model (given baseline knowledge of WaveNet, at least). The sample shown only does step 3 from your summary as far as I can tell.
In particular the top 2 samples from the Baidu blog most clearly show the gaps we need to cover from a research perspective to truly get "human level" TTS - a lot of the complexity in TTS is in the text part, and getting the subtleties of stress and f0 prediction correct is far from a solved problem. This is partly why there are so many different submodules and parts in the text piece of DeepVoice, and a whole appendix dedicated to that in WaveNet.
It is still surprising to me how good the audio models from WaveNet and DeepVoice are (such as the sample you show) - it seems that given good enough text features e.g. groundtruth the synthesis is nearly perfect. So it seems (IMO) that the next research papers will be focused on the text/f0 part.
I will also plug the work we have been doing at MILA, which is related but tries to directly go from text -> speech with attention based RNNs [0]. A longer version of our paper should be on arxiv soon, but for now we have a short teaser which was submitted as an ICLR workshop [1]. One fun feature is that we have the ability to handle multi-speaker synthesis, and multi-speaker datasets as well.
The primary differences from a high level are: we don't need a pronunciation dictionary for training, but they (DeepVoice or WaveNet) do. However we need a vocoder (currently WORLD) to get intermediate targets for training, while they do not. For training we need a directory of text files, and a directory of audio files, along with a script to run WORLD on the audio.
This means char2wav can go to new languages without extra lexical information (e.g. pronunciation dictionaries), since the WORLD assumptions are broad, but we find we have a difficult time on English without using phonemes (from a pronunciation dict), since the character -> audio mapping is pretty difficult for English.
We have better outputs (compared to char2wav English) on Spanish, German, and Romanian [2] though we are still improving every day. Note that the Romanian is still using fixed synthesis with WORLD - this is a demo only of the text -> WORLD features part of the model, the "reader".
The end result of both char2wav and DeepVoice (be able to go from text -> speech directly in generation) is the same. WaveNet also has this - you can see most of the TTS details in the appendix of the paper [3].
There is also a good discussion on reddit with one of the DeepVoice authors that has quite a bit of detail about their approach, I found it quite helpful [4].
Heiga Zen has a great overview talk on youtube for people who are interested in TTS [5].
There are also some interesting extensions on WaveNet [6][7], though I have not translated enough to determine what exact methods are being used.
[0] http://www.josesotelo.com/speechsynthesis/ http://www.josesotelo.com/speechsynthesis/
[1] https://openreview.net/pdf?id=B1VWyySKx https://openreview.net/pdf?id=B1VWyySKx
[2] https://www.youtube.com/watch?v=cwnDjq33uMs https://www.youtube.com/watch?v=cwnDjq33uMs
[3] https://arxiv.org/abs/1609.03499 https://arxiv.org/abs/1609.03499
[4] https://www.reddit.com/r/MachineLearning/comments/5wosbm/r_deep_voice_realtime_neural_texttospeech/ https://www.reddit.com/r/MachineLearning/comments/5wosbm/r_d...
[5] https://www.youtube.com/watch?v=nsrSrYtKkT8 https://www.youtube.com/watch?v=nsrSrYtKkT8
[6] https://twitter.com/ballforest/status/838759080621568002 https://twitter.com/ballforest/status/838759080621568002
[7] https://twitter.com/ballforest/status/836803448435789828 https://twitter.com/ballforest/status/836803448435789828