5 ms·
Ask HN: Any technical reasons Google Docs can't do voice typing in Firefox?
Google Docs enables voice typing when opened in Chrome. But not in Firefox.
Does somebody here know the technical reasons?
I find speech-to-text dictation a productive way to crank out a draft. So far, voice typing is the easiest approach on Linux that I've found. But I'd prefer Firefox. Any other suggestions for easy STT on Linux? Punctuation identification is preferred but optional. It's a Kubuntu system with only CPU, no GPU.
- FrenchDevRemote 4y agoAFAIK there is no technical reasons, just business reasons. Google systematically abuse their position to make Firefox less appealing to people
- mattzito 4y agoThere is absolutely a technical reason - Firefox doesn’t support the speech recognition API natively. https://developer.mozilla.org/en-US/docs/Web/API/SpeechRecognition https://developer.mozilla.org/en-US/docs/Web/API/SpeechRecog...
- FrenchDevRemote 4y agoSpeech recognition done locally on browsers is lame anyway, almost every site needing it use a custom solution.
- fxtentacle 4y ago"On some browsers, like Chrome, using Speech Recognition on a web page involves a server-based recognition engine. Your audio is sent to a web service for recognition processing, so it won't work offline." It's only about sending the speech to a server, anyway.
- SahAssar 4y agoIIRC every browser that supports the Web Speech API does so via cloud services. Mozilla being the only major browser engine maker without it's own "cloud" and having slightly fewer phone-home features than many of the others didn't want to do that. Mozilla has been doing quite a bit of work in the area though (for example https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech), hopefully to enable these features locally in the future. IMO it's good if we don't make web platform features unfeasible to implement locally, so I think mozillas stance makes sense. A new browser should not have to have the backing of a massive cloud provider willing to give away compute for free.
- fxtentacle 4y agoIt's purely business reasons. Also in Chrome, the STT runs in the Google cloud anyway. And just FYI I'm working on turning my state of the art speech recognition paper (TEVR, token entropy variance reduction) into a developer-friendly binary which will run offline on your GPU and, hence, be fully private, and offer a scripting API on localhost. Testing WER on LibriSpeech clean is 2.3, so slightly worse than an offline wav2vec2 1B, but years ahead of Kaldi, Vosk, coqui and the usual streaming cloud services. I estimate it'll be 2 more months until I can post Linux binaries. Won't be open source, though.
- lovelearning 4y agoPlease post a "Show HN" when you're done. I'd love to try it out!
- fxtentacle 4y agoHere's my Show HN for the open source CLI tool which is on GitHub with a MIT license: https://news.ycombinator.com/item?id=32409966 https://news.ycombinator.com/item?id=32409966
- lovelearning 4y agoThank you! I'll try it out this weekend.
- azmodeus 4y agoSounds awesome thanks for thinking about sharing your work!
- sillystuff 4y agoNerd dictation is a purely on-device speech to text program that works pretty well if your computer is fast enough. https://github.com/ideasman42/nerd-dictation https://github.com/ideasman42/nerd-dictation get speech models here: https://github.com/alphacep/vosk-api https://github.com/alphacep/vosk-api HN discussion: https://news.ycombinator.com/item?id=29972579 https://news.ycombinator.com/item?id=29972579
- lovelearning 4y agoThanks for the suggestions! I'll try them out.
- midislack 4y agoTechnically Google wants you on Chrome.
- b20000 4y agoyou can’t even do basic shit in a productive manner in google docs so what are you worrying about?