4 ms·
Can this be used to transcribe voice data in real time? I am building a docker image which will eventually accept in-browser audio via WebSockets outputting tr
by CommanderData 8y ago
Can this be used to transcribe voice data in real time?
I am building a docker image which will eventually accept in-browser audio via WebSockets outputting transcription in real time without needing Google WebSpeech.
https://github.com/ashwan1/django-deepspeech-server https://github.com/ashwan1/django-deepspeech-server
I planned to use DeepSpeech but this looks promising given it's low resources.
- themarkn 8y agoCould you tell me some more about this? I have used the web speech recognition API in Chrome (coming soon to Firefox also) to create a free project for real-time, editable, transcriptions in the browser that could be projected on a screen or subscribed to on a person's own device. I am frustrated by: - lack of cross-browser support on whatever device is generating the transcript. It will work from any Android phone but not from iOS - No offline functionality, because the work is actually not done locally - Poor accuracy.. we can correct this on the fly with the live editing, but it feels like it is not using the latest & greatest Google has to offer in terms of STT I don't know much about Docker or how to "use" a Docker image. Not asking you to teach me, but do you think your project would be useful for what I'm talking about? Currently using Firebase for the backend but it's really just exploratory at this point.
- iam-TJ 8y agoI too would like more information about both of your projects. I'm currently designing a digital technology platform for a charity for the blind in the UK and transcription of speech to text is one aspect I'm investigating, in addition to the more obvious text to speech.
- themarkn 8y agoHere's a video demo of what we are working on: https://youtu.be/xcUxd9sOkaM https://youtu.be/xcUxd9sOkaM There are a few features not listed (like exporting a correctly-formatted subtitle file) but you'll get the general idea. There's a link to an old demo in the video description.
- CommanderData 8y agoMy goal is to create a docker image with all necessary dependencies deployable with WebSocket endpoint exposed. Similar Watson S2T demo (https://speech-to-text-demo.mybluemix.net https://speech-to-text-demo.mybluemix.net) which utilizes WS. This is after I struggled to find anything in usable form on GitHub and Google Cloud pricing is prohibitive for my free projects. Perhaps there has been progress? web speech recognition API in Chrome I looked at this last year (See my comments 8-12 months back) and it seems either Google or Chromium dev team removed recognition.speechURI . It was there in Chrome however later removed. Would have been really helpful to just switch out provider from the browsers default to an alternative perhaps cheaper or free option. I don't know enough about the decision behind this, or if it was actually working in the first place. To be fair the whole "SpeechRecognition" is still in draft. See: https://developer.mozilla.org/en-US/docs/Web/API/SpeechRecognition/serviceURI https://developer.mozilla.org/en-US/docs/Web/API/SpeechRecog... End goal is to easily provide ASR/S2T through WebSockets. The uses and possibilities I have thought out just myself are many and I am sure it will be of use to others. > I don't know much about Docker or how to "use" a Docker image [....] but do you think your project would be useful for what I'm talking about? [....] Yes. But accuracy is probably something Google is winning on here. If your app is browser based, unless my assumption is incorrect you do not need an API key therefore are not bound by quota limits/costs.
- themarkn 8y agoThanks for expanding! Yes with the spec being in draft who knows what will change. I'm really interested in seeing what the Firefox implementation turns out to be. Hopefully close enough to Chrome's that everything still works. Are you on Twitter or is there some other way for me to know about your progress?