3 ms·
Here's the link to the paper: https://drive.google.com/file/d/0B3cxcnOkPx9AeWpLVXhkTDJINDQ/view https://drive.google.com/file/d/0B3cxcnOkPx9AeWpLVXhkTDJINDQ...
by e0m 10y ago
Here's the link to the paper: https://drive.google.com/file/d/0B3cxcnOkPx9AeWpLVXhkTDJINDQ/view https://drive.google.com/file/d/0B3cxcnOkPx9AeWpLVXhkTDJINDQ...
And the WaveNet site with audio samples: https://deepmind.com/blog/wavenet-generative-model-raw-audio/ https://deepmind.com/blog/wavenet-generative-model-raw-audio...
The comparison against state of the art Parametric and Concatenative methods are pretty mind blowing.
Particularly listen to the music samples. That's a generated piano piece that sounds quite musical.
They even include breaths and other auditory signals that really make for a convincing speech sample.
- exDM69 10y ago> Particularly listen to the music samples. That's a generated piano piece that sounds quite musical. That blew my mind. I've heard computer generated music before, but it's been just synthesized with usual methods while the computer is just the "composer". I find the crackling and buzzing a bit awkward, though. It's probably an artifact of the algorithm and can probably be mitigated with some simple filtering.
- deleted 10y ago[deleted]
- ChuckMcM 10y agoThanks for the links. It is definitely an improvement from the previous state of the art, although it seems surprisingly similar to artificial methods. And by that I mean that a number of RNN based network solutions for things like image classification or speech recognition produce output that is very unlike the non-RNN previous solution. This feels like they were training the system using artificially generated speech so that it could automatically generate artificially sounding speech.
- tlb 10y agoThe speech samples suffer from common flaws in prosody. When it says "The Blue Lagoon is a 1980...", it omits the pause after Lagoon. When giving a definition, there should be a comma-length pause after the term. The pacing would be correct for a non-defining sentence later in the same article like "The Blue Lagoon is a popular subject for parodies, including ..." Fixing that requires higher than sentence-level learning, since it has to know that it's introducing a potentially unfamiliar term which can only be known from the sentence's context within the article.
- jobu 10y agoThat is impressive. It seems like they're getting to the point where they can use a speech synthesis engine to train a speech recognition engine.