3 ms·
Is this really speech recognition from raw waveforms? It looks like they're extracting MFCC features from the raw audio, and using that as input to the neural n
by craigbaker 10y ago
Is this really speech recognition from raw waveforms? It looks like they're extracting MFCC features from the raw audio, and using that as input to the neural network. I thought that the point of WaveNet was that it took the raw waveform directly as input, unlike previous architectures which first extract spectral features such as MFCCs to use as the input.
- bmc7505 10y agoApparently, they tried to use the raw audio waveform with the original setup from the WaveNet paper but couldn't get it to train on their TitanX, so they used MFCCs instead. It's not exactly clear why this is the case. "Second, the Paper added a mean-pooling layer after the dilated convolution layer for down-sampling. We extracted MFCC from wav files and removed the final mean-pooling layer because the original setting was impossible to run on our TitanX GPU." [1] [1] https://github.com/buriburisuri/speech-to-text-wavenet#speech-to-text-wavenet--end-to-end-sentence-level-english-speech-recognition-using-deepminds-wavenet https://github.com/buriburisuri/speech-to-text-wavenet#speec...