3 ms·
While I appreciate the audio codec discussion and bandwidth-to-accuracy tradeoffs, how much of the speech recognition could be done on-device rather than shippi
by Amicius 7y ago
While I appreciate the audio codec discussion and bandwidth-to-accuracy tradeoffs, how much of the speech recognition could be done on-device rather than shipping it off to the cloud? It's my understanding that it's a matter of installing pattern files for analyzing the audio without needing to fail over to the cloud; how many GB are we talking to be able to cover normal daily speech, assuming a minimum of jargon? For the hearing impaired, not having to hit the cloud at all seems like the best option (and you don't need to compress the audio at all or worry about cloud-trip bandwidth).
- exikyut 7y agoOn Android: Language and Input > Google voice typing > Offline speech recognition, then ensure Wi-Fi and data are off and try shouting at textboxes (you might need to press a microphone button on your selected keyboard, unsure). It doesn't work very well in my experience.
- est31 7y ago> how many GB are we talking to be able to cover normal daily speech The models of Mozilla's DeepSpeech STT engine take 1.8 GB in compressed form: https://github.com/mozilla/DeepSpeech/releases/tag/v0.5.1 https://github.com/mozilla/DeepSpeech/releases/tag/v0.5.1 The main cause for the large size is the language model. They tried using different (smaller) language models, but they weren't as good.
- walterbell 7y agoGoogle published research [1][2] on offline recognition and it was rolled out earlier in 2019. Model size for English is claimed to be under 100MB, https://techcrunch.com/2019/03/12/googles-new-voice-recognition-system-works-instantly-and-offline-if-you-have-a-pixel/ https://techcrunch.com/2019/03/12/googles-new-voice-recognit... [1] https://arxiv.org/pdf/1603.03185.pdf https://arxiv.org/pdf/1603.03185.pdf (2016) [2] https://arxiv.org/pdf/1811.06621.pdf https://arxiv.org/pdf/1811.06621.pdf (2018)