3 ms·
I know it's difficult, but is there a way to gather feedback from the user when a voice command is below a certain threshold of understandability? Or to build a
by PennRobotics 3y ago
I know it's difficult, but is there a way to gather feedback from the user when a voice command is below a certain threshold of understandability? Or to build a context catalog, such as time of day specific types of commands are expected?
I think these two areas could be holding the commercial voice assistants back, since the majority of bug reports on the Google Home and Alexa subreddits are people complaining about being misheard.
- kkielhofner 3y agoI've been thinking about this a little recently. Raw brain dump incoming... The problem with intent matching now (say with Home Assistant) is that the transcribed speech more-or-less needs to exactly (character for character) match the defined entities (names of lights, etc) as well as the grammar and structure expected by the intent definitions. Something as simple as a hyphen in "turn-off" currently breaks the matching. That's an easy one to fix, I'm just providing it as an example. In terms of addressing this, I've kicked around a few ideas: With Home Assistant at least we can pull all of the entities and supported intents. We have a variety of options to do all kinds of fuzzy matching with these now known phrases. We'd fix the transcription based on something like nearest neighbor (more or less) and send the clean command to HA. There is also a way to handle this in the prompt to Whisper. A literal text embedding model with nearest neighbor search seems kind of ridiculous but I've been curious about it. The same could be done at the audio level by essentially making an embedding of the Mel Spectrogram of the actual audio (which we could have available) and searching on that. There are also approaches with a variety of NLP implementations, language models, etc that could be combined or used separately. We've also considered some more manual approaches - things like an interface that logs sessions where you can (after the fact) find a session/command that went wrong, and basically map it to what you wanted. While this may seem tedious we get some reports that our speech recognition consistently mis-transcribes a given voice command. Implementing a lookup table to correct for this is simple. The biggest challenge to this is how to deal with the UI/UX aspects.