3 ms·
On "...clearly wouldn't be used for real speech synthesis" - I think it depends what language you are looking at. For languages with many years of linguistic re
by kastnerkyle 10y ago
On "...clearly wouldn't be used for real speech synthesis" - I think it depends what language you are looking at. For languages with many years of linguistic research (e.g. English), it will be hard to beat a good parametric TTS or even a well-engineered concatenative system that encodes years and years of linguistic knowledge/features, and also has the ability for engineers to add pronunciation rules for words which are wrong. But for languages which are not as focused on, there are some gains to be made in my opinion.
See some early work on Romanian (https://www.youtube.com/watch?v=cwnDjq33uMs https://www.youtube.com/watch?v=cwnDjq33uMs), and compare to Google Translate TTS (the second to last, robotic example). The existing TTS systems out there in the research community are also pretty good (last example), but I think we are at least competitive which is interesting, given that we are far from TTS experts especially in all the NLP processing that normally happens in building these systems (https://github.com/CSTR-Edinburgh/merlin/blob/master/misc/questions/questions-radio_dnn_416.hed https://github.com/CSTR-Edinburgh/merlin/blob/master/misc/qu... for example).
At very least the Edinburgh system (http://romaniantts.com/new/ http://romaniantts.com/new/) should be the baseline across companies for Romanian - but one of the things we also want to show is that one architecture/approach can generalize easily across several languages. By and large the approach for all our languages is identical, including nearly all hyperparameters (we change one setting for attention default step size, and I think that is it).
There are tons of languages which have poor existing systems, and having something that is basically - record a speaker for a while, write down whatever sentences they are saying, train a big model, and do pretty well - could be very useful. We are still exploring what languages this approach works for, but have had a good success rate so far if we can find a good openly available dataset. It also opens the door to using existing sentence level ASR datasets "in reverse", since we don't need timing/alignment information.
There is also a lot of potential for personalized sound/speakers using speaker interpolation (see the alternate speaker examples in the youtube video) that we have not explored yet, as well as applications to related sequence generation tasks. I think some of our training tricks are useful for training sequence generation models in general.
Some great videos that help put our work in context [1][2].
[0] early Romanian Demo of our approach: https://www.youtube.com/watch?v=cwnDjq33uMs https://www.youtube.com/watch?v=cwnDjq33uMs
[1] Alex Graves, Generating Sequences with Recurrent Neural Networks (see ~36m in for speech demo): https://www.youtube.com/watch?v=-yX1SYeDHbg https://www.youtube.com/watch?v=-yX1SYeDHbg
[2] Heiga Zen, Generative Model-Based TTS Synthesis: https://www.youtube.com/watch?v=nsrSrYtKkT8 https://www.youtube.com/watch?v=nsrSrYtKkT8