5 ms·
Mozilla DeepSpeech trained on the Common Voice dataset for English. You can get pretrained models too. They have a nice matrix channel where you can get help, a
by ftyers 7y ago
Mozilla DeepSpeech trained on the Common Voice dataset for English. You can get pretrained models too. They have a nice matrix channel where you can get help, and pretty good documentation. It is also actively developed by several engineers. http://voice.mozilla.org/en/datasets http://voice.mozilla.org/en/datasets and http://github.com/mozilla/DeepSpeech/ http://github.com/mozilla/DeepSpeech/
- dmos62 7y agoThe models (`deepspeech-0.6.1-models.tar.gz`) weigh 1.14 gb, if anyone's interested.
- kingo55 7y agoInteresting. Are there any projects actively using this today?
- nmstoker 7y agoI haven't tried it directly myself but there is this project, Dragonfire, which looks quite reasonable using DeepSpeech: https://github.com/DragonComputer/Dragonfire https://github.com/DragonComputer/Dragonfire There's a minimal demo app I put together here too: https://github.com/nmstoker/SimpleSpeechLoop https://github.com/nmstoker/SimpleSpeechLoop
- dsteinman 7y agoI've been using DeepSpeech to learn to build voice controls for all sorts of things in JavaScript. And I've got a way to connect it to the web so you'll be able write speech recognition enabled web pages using client side JavaScript. https://github.com/jaxcore/deepspeech-plugin https://github.com/jaxcore/deepspeech-plugin
- magicalhippo 7y agoI had limited luck using the provided language model, but very good results when providing my own. So if you feel the results are poor, try building your own language model. AFAIK DeepSpeech works by using the neural net to detect characters from speech, and then the language model is used to try to make a sentence out of the character stream, by doing a kind of graph search. Thus if the language model doesn't contain the words you want it to recognize, it'll have a hard time giving good output. Anyway, I used the following tutorial[1] as a base to build the language model. For the kenlm tools I used Ubuntu WSL, and the generate_trie executable was part of the DeepSpeech native tools package for Windows. [1]: https://discourse.mozilla.org/t/tutorial-how-i-trained-a-specific-french-model-to-control-my-robot/22830/2 https://discourse.mozilla.org/t/tutorial-how-i-trained-a-spe...
- ghostpepper 7y agoFor a software engineer with no experience in machine learning / AI, what does it mean to build your own language model? Does it require coding? Hundreds of hours of audio data from your own voice? A significant amount of computing power?
- magicalhippo 7y agoThe tools available means you only need to provide a list of normal sentences, and they should include the words you'd like it to know about. For my case I just wanted to train it on like 30 different sentences, that took less than a second. But for a general assistant ala Google Home you'll want a large number of sentences and I hear it can take a while (hour or few?). Due to using probabilities it will match words in other sentences than what you give it, but from my understanding it will be partial to the ones you feed it if DeepSpeech mis-classifies a character or two.
- nmstoker 7y agoHere's an example I did using a custom LM with DeepSpeech - the description links back to the forum with the steps for producing it. https://youtu.be/LWUBK6PAaxM https://youtu.be/LWUBK6PAaxM This was on a slightly earlier version, and they've made improvements in speed and quality of recognition since then.
- magicalhippo 7y ago> Hundreds of hours of audio data from your own voice? I should clarify this. As I mentioned, training the neural net part requires tons of audio and the corresponding text (and people should totally contribute[1], the resulting data sets are released to the public). The neural net in DeepSpeech is then used on an audio stream and outputs a stream of characters. Turning that stream of characters into sentences is what the language model is for. Training the neural net is very data and compute intensive, but fortunately Mozilla provides pre-trained models. Generating the language model is relatively cheap. And if your target language shares sounds with English, you may get away with using the English-trained neural net but with a non-English language model. [1]: https://voice.mozilla.org/ https://voice.mozilla.org/
- hutzlibu 7y agoHere is a paper from a german university, where they tried to adopt DeepSpeech to german: https://www.researchgate.net/publication/336532830_German_End-to-end_Speech_Recognition_based_on_DeepSpeech https://www.researchgate.net/publication/336532830_German_En... The results are, that adopting works good, but as of writing of the paper somemonths ago, the results were not very good yet. So it takes better trained models for other languages. English seems to be quite good. What surprised me, is that it works offline very fast even on a rasperry pi! https://www.hackster.io/dmitrywat/offline-speech-recognition-on-raspberry-pi-4-with-respeaker-c537e7 https://www.hackster.io/dmitrywat/offline-speech-recognition...