3 ms·
Thanks! In tests with a high quality headset mic, I've been able to get up to 98% (intent) recognition accuracy. Keep in mind that this is for a closed domain w
by synesthesiam 7y ago
Thanks! In tests with a high quality headset mic, I've been able to get up to 98% (intent) recognition accuracy. Keep in mind that this is for a closed domain with only ~600k possible sentences. I'd expect performance degradation with a free-standing mic like the PS3 eye, but I don't have any measurements yet.
Because Rhasspy combines both a speech and intent recognizer, transcription word error rate isn't as important if the correct intent/named entities can still be recognized. When your intents have more distinct vocabularies, your recognition rate will be higher.
I've had folks with British and Australian accents use Rhasspy successfully (with the CMU English 5.2 acoustic model). Some recent tests with an nnet3 Kaldi model have produced better results (https://github.com/gooofy/zamia-speech https://github.com/gooofy/zamia-speech), especially with noise. I plan to add this model to Rhasspy in the near future.
If you're specifically thinking of an Indian accent, I noticed that CMU recently published an Indian English acoustic model (https://sourceforge.net/projects/cmusphinx/files/Acoustic%20and%20Language%20Models/Indian%20English/ https://sourceforge.net/projects/cmusphinx/files/Acoustic%20...). I'd be happy to add this to Rhasspy if you're interested.