4 ms·
You might be interested in trying Rhasspy: https://github.com/synesthesiam/rhasspy https://github.com/synesthesiam/rhasspy Rhasspy lets you describe the set of
by synesthesiam 7y ago
You might be interested in trying Rhasspy: https://github.com/synesthesiam/rhasspy https://github.com/synesthesiam/rhasspy
Rhasspy lets you describe the set of sentences you want to speak using a simple grammar with annotations for named entities (https://rhasspy.readthedocs.io/en/latest/training/#sentencesini https://rhasspy.readthedocs.io/en/latest/training/#sentences...). It outputs JSON over HTTP/Websockets/MQTT, so it works well with NodeRED, Home Assistant, etc.
Disclaimer: I created and maintain Rhasspy.
- nmstoker 7y agoThat looks really impressive, integrating a number of open source tools with a simple but nice UI. Do you have any measures of how well it recognises spoken commands? And have you seen anyone using it with non-American accents for English? (I ask as it relies on the CMU dictionary and tools I've seen use it tend to struggle with other accents, understandably)
- synesthesiam 7y agoThanks! In tests with a high quality headset mic, I've been able to get up to 98% (intent) recognition accuracy. Keep in mind that this is for a closed domain with only ~600k possible sentences. I'd expect performance degradation with a free-standing mic like the PS3 eye, but I don't have any measurements yet. Because Rhasspy combines both a speech and intent recognizer, transcription word error rate isn't as important if the correct intent/named entities can still be recognized. When your intents have more distinct vocabularies, your recognition rate will be higher. I've had folks with British and Australian accents use Rhasspy successfully (with the CMU English 5.2 acoustic model). Some recent tests with an nnet3 Kaldi model have produced better results (https://github.com/gooofy/zamia-speech https://github.com/gooofy/zamia-speech), especially with noise. I plan to add this model to Rhasspy in the near future. If you're specifically thinking of an Indian accent, I noticed that CMU recently published an Indian English acoustic model (https://sourceforge.net/projects/cmusphinx/files/Acoustic%20and%20Language%20Models/Indian%20English/ https://sourceforge.net/projects/cmusphinx/files/Acoustic%20...). I'd be happy to add this to Rhasspy if you're interested.