3 ms·
Bottom line: speech recognition in the general case (more than a few predetermined words) is only as good as the 1) acoustic model (which utterances were heard)
by plainsman 14y ago
Bottom line: speech recognition in the general case (more than a few predetermined words) is only as good as the 1) acoustic model (which utterances were heard), and 2) language model (how do we group the utterances into words).
This requires massive amounts of labeled data. This is why Nuance is king and few others come close - the amount of labeled data necessary to catch up is astounding. Not to mention a patent minefield to navigate.
This is unfortunately one field in which open-source alternatives face real obstacles and won't be viable in the near future.
- DennisP 14y agoSeems like there ought to be a way to crowdsource some of that.
- 3amOpsGuy 14y agoYeah there is http://www.voxforge.org/home http://www.voxforge.org/home I have CMU Sphinx (pocket edition even though I'm running on a full blown server) and for my use case it works fairly well.
- plainsman 14y agoThis cool. But are these audio files transcribed, or just provided? I downloaded a couple here: http://www.repository.voxforge1.org/downloads/SpeechCorpus/Trunk/Audio/Main/16kHz_16bit/ http://www.repository.voxforge1.org/downloads/SpeechCorpus/T... and it didn't seem to have a log of words labeled each by timestamp offset into the audio recording - which is the vital part for training a recognizer. Am I missing something?
- 3amOpsGuy 14y agoIt's trickier than just matching word sounds, the sphinx docs are first rate: http://cmusphinx.sourceforge.net/wiki/tutorialam http://cmusphinx.sourceforge.net/wiki/tutorialam It is very interesting but unfortunately just appears to be too much hassle for sane people to tackle (although it'd be extremely worthwhile if someone would innovate in this space and lower the barrier to entry - most are using commercial acoustic models with the FOSS software)