4 ms·
Hybrid DNN-HMM type systems (or CNN-HMM) still roundly beat purely RNN based systems on public benchmarks (~5% across tasks such as Switchboard, WSJ, even a bit
by kastnerkyle 10y ago
Hybrid DNN-HMM type systems (or CNN-HMM) still roundly beat purely RNN based systems on public benchmarks (~5% across tasks such as Switchboard, WSJ, even a bit behind on TIMIT) and most of the best internal systems at the companies I know (barring Baidu) are still using this DNN-HMM scheme, though RNN+CTC or CNN+CTC is not far behind performance wise these days.
RNN or CNN+CTC is tantalizing because it is conceptually simple, should "scale with data", and you can build a toy version in a few hundred lines of code! Which is a far cry from the massive complexity of most speech systems in use today.
I think we are still a ways off from replacing all HMMs with RNNs, but with tighter integration of the language model into the speech system and training the whole shebang end-to-end there will be some interesting results. There are a few ares of machine translation where people have tried this "deep fusion" with promising improvements.