5 ms·
A 2019 Guide to Speech Synthesis with Deep Learning
- juris-ws 7y agoWe're experimenting with this in combination with deepfake videos. https://wiserstate.com https://wiserstate.com Spooky to think that one day we might be able to digitally "clone" ourselves this way.
- aswanson 7y agoWell written. Makes me want to open a medium account and explain something to make sure I'm not getting rusty.
- mwitiderrick 7y agoYou should give it a shot
- aswanson 7y agoThanks for the encouragement. And the great article.
- ghaff 7y agoOr your own blog. Though, honestly, for very occasional stuff, Medium probably makes more sense.
- mgradowski 7y agoPardon the interruption, a free static page hosting would respect the reader a little bit more than Medium.
- ghaff 7y agoPersonally I would (do) go with a free hosted service like Blogger which will handle big traffic spikes. But if someone is just wanting to push out a blog post or two a year, I'd be hard put to argue against Medium.
- chrisa 7y agoIf you're looking to write coding related content - I recommend dev.to as well! It has a lot of the same benefits as medium (built in readership), without all the popups, etc :)
- PieSquared 7y agoI'm an author on a few of these papers referenced (the Deep Voice papers from Baidu). I'm happy to answer any questions folks may have about neural speech synthesis, as I've been working on this for several years now. In general, it's a fascinating space. There are challenges in text processing (not even mentioned in the blog), such as grapheme to phoneme conversion, part of speech detection, word sense disambiguation, text normalization, challenges in utterance-level modeling (spectrograms), and challenges in "spectrogram inversion" / waveform synthesis. The NLP components of the pipeline are often overlooked but are no less important than they were a few years ago -- part of speech / word sense is the difference between "Time is a CONstruct" and "I'm going to conSTRUCT a tower", and is the difference between "Let's drop that bass" being about a DJ or about a fish. The acoustic modeling phase (e.g. Tacotron, Deep Voice 3) works fairly well, and can produce some awesome demos with things like style tokens ("GST-Tacotron"), but still has a ways to go until it can encompass the full range of human inflection and emotion. At the waveform synthesis level, models like WaveRNN (with subscale modeling) and Parallel WaveNet make it possible to deploy modern waveform synthesis models, but it's still a major issue to deploy them onto low-power devices due to compute restrictions. Overall, lots of interesting challenges to work on, and we're making a lot of progress quite quickly; and I haven't even started talking about voice conversion or voice cloning!
- mwitiderrick 7y agoGreat work on the paper!
- nfoz 7y agoCool! Does text-to-speech require AI, or is there any active work in non-AI methods? Which bits are the AI bits? Do "deep" methods substantially improve over whatever classical methods we might have had?
- PieSquared 7y agoI'll try to answer these one at a time. 1. Does text-to-speech require AI? This one is a bit tricky to answer since it requires defining "AI". AI as a moniker has been used to describe deep neural networks, search algorithms, expert systems and logic systems, particle filters, SVMs, etc etc. Almost all text-to-speech (TTS) systems are based on a combination of some of these machine-learning methods and digital signal processing (DPS), so I would say yeah, text-to-speech is exactly what AI describes, even if it doesn't resemble human-like thinking like other AI applications do. 2. Is there any active work in non-AI methods? This one again is a bit tricky for the same reason as before. However, there's a ton of pieces of the TTS pipeline that aren't AI in the current sense of the word (machine learning with neural networks or HMMs or other classifiers). For example, concatenative systems will traditionally take a large database of audio, divide it into chunks, and then recombine those chunks, using some interpolation method such as (OLA, PSOLA) to overlap those chunks. Choosing the chunks to overlap to create the target speech becomes an AI / search problem, using some sort of acoustic model to predict the acoustic parameters of each frame and then using a Viterbi search algorithm with target / join costs to find the optimal chunks. As another example of non-AI parts of the pipeline, text normalization tends to involve a lot of hand-written rules; for example, should you say "5/10/2019" as "May tenth, twenty nineteen", "the tenth of may twenty nineteen", "the tenth of may two thousand nineteen", or even "october fifth twenty nineteen". This decision and the conversion is often done with a ton of handwritten rules or grammars (see Kestrel, Google's text normalization system, and the open-source version, cleverly named Sparrowhawk). Anyways, the real answer is that TTS is always a combination of AI (machine learning) approaches with specialized text and audio processing algorithms. 3. Which bits are the AI bits? The AI bits are the bits where you need to make some sort of heuristic decisions, and you'd like to make them by imitating some target speech. For example, things like part of speech detection, predicting acoustic parameters (spectrograms, F0, etc), more recently waveform synthesis as well. 4. Do deep methods significantly improve on the state of the art? Yes, though they also come at a cost. For example, deep sequence-to-sequence networks make great frame-level models: Tacotron and similar models can do things like emotional and stylized voice synthesis much better than what I've seen HMMs and other non-deep models do. As another example, WaveNet / WaveRNN / etc are some of the only parametric speech models (that is, generating the waveform from scratch instead of copying it from a database of audio) that can match the quality of concatenative models (copying audio from a database), but they can be quite difficult to deploy due to high computational cost. Overall, though, yeah, deep methods and all the improvements to neural networks in the past few years are having a profound impact on the quality and naturalness of TTS.
- amelius 7y agoRepeated typo. "Casual convolutions" should probably be "causal convolutions".
- mwitiderrick 7y agoThanks for the noticing that.