5 ms·
This model is quite cool, but also quite a bit different than what lyrebird.ai is doing. NPSS has a lot of extra information in the control inputs about pronunc
by kastnerkyle 9y ago
This model is quite cool, but also quite a bit different than what lyrebird.ai is doing. NPSS has a lot of extra information in the control inputs about pronunciation and timing (the part-of-phoneme timer feature) - this means that most of the "hard parts" (in my opinion) for naturalness are control inputs to NPSS/WaveNet style models, rather than variables the model must generate globally and consistently as in lyrebird. At generation time NPSS appears to generate each component autoregressively as well, but I am not clear on whether the demo samples do this or if they use "true" values for f0 at least - what forces the model to sing the exact same melody, if many melodies are possible given the underlying audio information?
Also note that NPSS has some amount of post-processing, at least reverb and perhaps other common musical mixing - we don't really know how these samples are generated, and I have a hard time decyphering exactly what inputs are required, and what are generated from the paper alone. However, I really, really, really like NPSS - I just don't think the comparison you are making is valid here.
These features (f0, duration, pronunciation) are some of the most difficult things to learn to model from datasets of speech and text directly, and I am not sure how they got the subset used (I think only f0 and pronunciation/phoneme) for this NPSS model. Giving creators fine-grained control of the performance (as in NPSS) is quite cool, and if these systems can get fast enough I think the possibilities are really exciting. The same things could likely be done with lyrebird as well - there is no real "tech reason" you couldn't add more conditional inputs, with finer grained information/control.
The key part in my mind is deciding what amount of complexity to show to a user, and what amount to try and capture inside the model - some people may want to control (for example) duration and f0 directly for a performance, while others may want to just upload clips to an API and get reasonable results back, with less ability to control each sample (they can still curate themselves for the "best" samples). Lyrebird.ai is handling the latter case, while the former case would require quite a bit more intervention from the average user, almost becoming like an instrument ala the original voder [0]. However, you could potentially have both approaches as a kind of beginner/advanced mode, but advanced mode needs a user interface, and probably near-realtime feedback.
I used to really strongly believe that the audio model was going to be the hard part of "neural" TTS (blame my background in DSP perhaps), but post-WaveNet the game has really changed a lot - conditional audio models are something we are starting to know how to do pretty well.
The text pipeline of most TTS systems is still the craziest part in my mind, check out a "normal" feature extraction of 416 hand-specified features [1]! These extractions can be upwards of 1k features per timestep/frame, and generally require a lot of linguistic knowledge to specify for new languages. It seems (given Alex Graves' demo [2], char2wav [3], tacotron[4]) that we are making progress on learning this information directly from text, which in my mind is a key breakthrough for TTS in languages besides English, where lots of work on English pronunciation has been done already and is generally available.
[0] https://www.youtube.com/watch?v=TsdOej_nC1M https://www.youtube.com/watch?v=TsdOej_nC1M
[1] https://github.com/CSTR-Edinburgh/merlin/blob/master/misc/questions/questions-radio_dnn_416.hed https://github.com/CSTR-Edinburgh/merlin/blob/master/misc/qu...
[2] https://www.youtube.com/watch?v=-yX1SYeDHbg&t=38m00s https://www.youtube.com/watch?v=-yX1SYeDHbg&t=38m00s
[3] http://josesotelo.com/speechsynthesis/ http://josesotelo.com/speechsynthesis/
[4] https://google.github.io/tacotron/ https://google.github.io/tacotron/
- yishhh 9y agoHi Kyle, I was wondering if the lyrebird github implementation will be open sourced as currently I am hoping to work on improving the current implementation by incorporating prosody into speech synthesis, thanks!