3 ms·
Still based on MFCC features, which are very noisy / high entropy.
by lobius 10y ago
Still based on MFCC features, which are very noisy / high entropy.
- gok 10y agoThis is completely wrong. Kaldi has code to compute all kinds of features (filter bank, pitch, PLP...) and the many recipes use them.
- lobius 10y agoBut not spectral peaks, which is what audio fingerprinting services like Shazam use (very successfully too it seems)
- skoocda 10y agoThere's been very little use of spectral peaks in LVCSR for the past ~10 years.
- unlikelymordant 10y agoShazam and asr are very different problems. The things that work for one don't work for the other.