3 ms·
Pretty much SOA; most end-to-end systems use these spectrograms on short time slices. The alternative is mel-frequency cepstral coefficients, which are used mor
by skoocda 10y ago
Pretty much SOA; most end-to-end systems use these spectrograms on short time slices. The alternative is mel-frequency cepstral coefficients, which are used more in GMM-HMM speech recognition than for DNN.
- annnnd 10y agoFor other illiterates like me: * SOA = state of the art * GMM-HMM = Gaussian mixture model - Hidden markov model * DNN = deep neural networks (more than 1 hidden layer) * mel-frequency cepstral = MFC ;-)
- kwhitefoot 10y agoCould sum1 write a GM script to replace all abbr.s with their ffs? The over use of abbr.s is one of the most XABs on Hacker News. Could someone write a Grease Monkey script to replace all abbreviations with their full forms? The over use of abbreviations is one of the most extremely annoying behaviours on Hacker News. (Thank you https://www.allacronyms.com/aa-search?q=annoying&cx=010821384832661523411:tvncv4ludxy&cof=FORID:11 https://www.allacronyms.com/aa-search?q=annoying&cx=01082138...)
- zump 10y agoHow does this work when the spectrogram is finite time slice?!
- skoocda 10y agoLots of overlapping. It's a sliding window function. Ballpark for most algorithms: 10 ms of new audio, 90 ms of old audio.
- eutectic 10y agoWhy not just use time-domain convolutions?