3 ms·
Why don't you use STFT + Conv2D like Deep Speech 2 did. It works well in my case.
by stealthcat 8y ago
Why don't you use STFT + Conv2D like Deep Speech 2 did.
It works well in my case.
- p1esk 8y agoThe DeepSpeech2 paper does not include any details about audio processing. I see an older Baidu-Research implementation of DS1 that uses "log of linear spectrogram from FFT energy". Also, there's a pytorch implementation [1], where they use Librosa's STFT, is that what you're referring to? That's two more implementations that I haven't considered. I'm sure most of the processing steps under the hood are the same or similar, but as I'm not an audio processing expert, I can't tell which method is better (and why). And it's hard to tell if it "works well" because or despite the way I processed the files. [1] https://github.com/SeanNaren/deepspeech.pytorch https://github.com/SeanNaren/deepspeech.pytorch