3 ms·
Long 5am post alert broken into sections corresponding to the authors. I wish I read this 8 months ago so it could of prepared me for some of the misery to com
by Caligula 17y ago
Long 5am post alert broken into sections corresponding to the authors.
I wish I read this 8 months ago so it could of prepared me for some of the misery to come. I initially thought it would be simple. Use some of the open source tools or commercial SDK's, then quickly move on to the important stuff. Did not turn out that way. In fact, speech recognition is a time killing whore, but it sure is interesting, sometimes at least.
1. Telephony & VoIP:
I disagree with the author. Freeswitch comes with speech recognition built in and recently added uniMRCP so I think FS is definitely better not even taking into account how shady asterisk is.
The prices he lists are also on the high side. You can get local DID's for half what he lists, probably further less in bulk.
2. Web Services:
Agree with his comments on web services. Speech recognition is perfect for scaling onto the magical cloud. Dictation is more difficult because for very large vocabulary at least, your going to take up a system to process it and be lucky to get 1RT.
3. Embedded
Disagree with this. Some decoders are made specificially for this. Well pocketsphinx is. It even used ARM ASM code to speed up calculations on embedded devices so you can get better results on the iphone for example. The developer for pocketsphinx for example recently made this fantastic demo for the N800 which I recall was similar to the iphone in power.
http://www.youtube.com/watch?v=OEUeJb6Pwt4 http://www.youtube.com/watch?v=OEUeJb6Pwt4
And I am sure it can be tweaked to be even better. So much of speech recognition is tweaking. As he states correctly, Lumenvox is based on sphinx(sphinx2 I recall reading), just its very tweaked.
Commercially available:
Very funny, I agree with most. MS,AT&T,IBM suck at least for providing API's but their tech is very good. IBM for example released theirs as opensource but changed their mind and just let it die. Nuance is 'contact our salesperson' expensive. Lumenvox is the most affordable.
Open Source:
Minor nitpick, he should of included HTK with Julius. Acoustic models definitely have given me the most grief.
I use sphinx4 and pocketsphinx and am very pleased. They are state of the art decoders. The acoustic model, or lack thereof is the reason why commercial engines are perceived as superior. If you have 50k you can spend get a LDC membership with tonnes of transcribed data. Slap it into the trainers format, BAM. Unfortunately I was not going to do this. I disliked the LDC for what I thought was gauging. Same prices for mega corporations as individuals. But after making my own model, and still in the neverending process of making my own, tweaking it, etc.., I appreciate the misery that is collecting and organizing transcribed data and appreciate the work they do even if I can't use it.
Notice something wrong with your model, need to retrain it. Takes days with a quad core. In fact my new 8core beast of a server I ordered arrives next week, I wonder how my parents will take the noise, apparently servers are loud. Even with that it will take maybe half a day. And spotting errors is hard. I cant emphasize enough how boring it is to listen to hours on end of audio and see that it matches up with the text perfectly. In some cases, listening a bunch of times to make sure. Noticing issues with your model, having to go figure out why.
Voxforge.org is great. They have ~50 hours of quality data, most at 16khz computer microphone but a good portion at 8khz telephone. You can always downsample. But for dictation you need much more.
You dont need thousands of hours unless your doing dictation and if that extra few percent is worth it. You can get good results with low hundreds. There are other equally important factors like language models that he should of mentioned that could be equally as important as the acoustic model. How its important to have relevant, and lots of data to train them. The acoustic model is only one of many factors(as is the decoder for the matter). Perhaps because its that he did not focus on dictation that he left it out.
Wrapping Up:
FS better. At least try both, its trivial to set each up and follow a simple tutorial. You can still plug lumenvox into FS, although it will cost. But it would cost the same for asterisk. I agree that its difficult but I don't think that should stop you. Just be aware its a lot of work, some of it very boring and frustrating, but I am sure the same can be said for most things. Ok maybe not :)