4 ms·
How accurate is CMU Sphinx for speech recognition compared to what's inside Alexa?
by tianlins 10y ago
How accurate is CMU Sphinx for speech recognition compared to what's inside Alexa?
- IshKebab 10y agoSphinx is pretty awful (remember the time before good speech recognition existed?). Alexa is far better. Kaldi is much better, but very difficult to set up. None of the open source speech recognition systems (or commercial for that matter) come close to Google.
- dharma1 10y agohttps://github.com/alumae/kaldi-gstreamer-server https://github.com/alumae/kaldi-gstreamer-server Takes the pain out of it
- IshKebab 10y agoAh interesting link, I hadn't seen that.
- amelius 10y ago> None of the open source speech recognition systems (or commercial for that matter) come close to Google. Is that because of the data they have, or because of their superior algorithms?
- IshKebab 10y agoBoth I think, but mostly the data. Baidu's deep speech is meant to be very good and its design is public (they even open sourced one component of it).
- kuschku 10y agoIt’s because they used data from the public to train their models. If, suddenly, someone would apply the fact that copyright bans remixes to training of neural networks, and apply the fact that licenses for this have to be granted explicitly, Google would lose 90% of their advantage over other companies. Personally, I’d be for making a requirement that companies open source their trained models if the training data contained data supplied by users, not paid employees.
- amelius 10y agoI sympathize, but good luck with that :)
- davexunit 10y agoI think the world needs the equivalent of OpenStreetMap but for speech data, so that the data is under a copyleft license that legally enforces reciprocation when the corpus is used or modified.
- ashitlerferad 10y agoThe closest is VoxForge: http://voxforge.org/ http://voxforge.org/
- davexunit 10y agoDidn't know about VoxForge. Thanks!