7 ms·
A 2019 Guide for Automatic Speech Recognition
- konz 7y agoRelated guide from last week: A 2019 Guide to Speech Synthesis with Deep Learning (https://news.ycombinator.com/item?id=20819672 https://news.ycombinator.com/item?id=20819672)
- m4tthumphrey 7y agoOff topic: Why does that site not show the scrollbar? I use it to work out how long a post is...
- skykooler 7y agoSo how can one actually use one of these systems? I'm familiar with pocketsphinx, where you download it and then run procketsphinx_continuous and it prints a transcription of the microphone. Or Julius, where you write a grammar and run it with that and it prints the output. The process for DeepSpeech seems to be along the lines of "get access to a machine with ten GPUs, find a huge dataset, train it, and then run that somehow" - are there compiled models that can actually be shipped as part of an application?
- nwalker85 7y agoVerbio offers an ASR with built-in grammars, which is what you are asking for. Deepspeech is hot garbage in my experience.
- jamesonthecrow 7y agoObviously the big cloud players offer their own APIs and SDKs (for a price), but there are a few other solutions worth looking at. Facebook has open sourced some pre-trained models: https://github.com/facebookresearch/wav2letter https://github.com/facebookresearch/wav2letter Picovoice has some smaller, more efficient models capable of running on edge devices: https://github.com/Picovoice https://github.com/Picovoice Full ASR does require quite large models and datasets, but you don't need nearly that much power or data to fine-tune a model for your own domain.
- skykooler 7y agoAh, thanks, I wasn't aware Picovoice had released an actual engine yet.
- GordonS 7y agoWasn't aware of Picovoice. Just tried the live do they have on their website... wow, it's... not great! Even if I spoke as precisely as possible or/and put on an American accent, it was way off the mark.
- lunixbochs 7y agoI made a realtime mac demo of the acoustic model part of wav2letter++ including several trained models: https://github.com/facebookresearch/wav2letter/issues/327 https://github.com/facebookresearch/wav2letter/issues/327 Someone in that thread ported it to linux. This demo is just "acoustic model emissions", which are character level predictions (with no repeated characters), but I have "decoding" (turning into english sentences) working locally as well and I'll post a new demo at some point.
- TaylorAlexander 7y agoMozilla is maintaining an implementation of Baidu's deepspeech with pretrained models (and checkpoints so you can fine tune with your own dataset). They have an example that accepts streaming from the microphone: https://github.com/mozilla/DeepSpeech/tree/master/examples/mic_vad_streaming https://github.com/mozilla/DeepSpeech/tree/master/examples/m... See the last full release here: https://github.com/mozilla/DeepSpeech/releases/tag/v0.5.1 https://github.com/mozilla/DeepSpeech/releases/tag/v0.5.1
- gok 7y agoI'm a little confused about the title because the first paper is from 2014. It's also too bad this doesn't mention any traditional HMM-based ASR techniques, as HMMs continue to be used on many SOTA systems, particularly those that can be reproduced publicly: https://github.com/syhw/wer_are_we https://github.com/syhw/wer_are_we
- priansh 7y agoThis. The article quotes DeepSpeech, wavenets, LSTMs of all sorts; essentially, all the neural networks that scale terribly. DeepSpeech for example is pretty heavy and requires a decent GPU to get anywhere near a realtime factor of 1. Meanwhile ASR through HMM's consistently hits realtime factors sub-1 and can run on small CPUs. e.g. the default models Kaldi ships with outperform DeepSpeech on a lot of modern examples, and are exponentially faster. Moreover, training these HMM's is something that is feasible for a normal developer. Training newer models requires data of scale and quality (iirc Mozilla's models are trained on Common Speech which is an enormous crowd sourced dataset, and Google's wavenet models use an internal dataset of very high quality and quantity). Until the models get more practically achievable, ASR for average people will probably continue to be dominated by Kaldi, Sphinx etc
- GordonS 7y agoDo you have any advice or links for a "normal" developer to get started with HMMs for speech recognition? I'd love to build something just for myself here.
- priansh 7y agohttps://github.com/jcsilva/docker-kaldi-gstreamer-server https://github.com/jcsilva/docker-kaldi-gstreamer-server ^^ If you're getting started, just following the steps here can get you set up really fast. IIRC there's also an HTTP endpoint at /recognize you can use instead of WebSocket if you're transcribing audio files so it's pretty cool!
- 7y ago
- GordonS 7y agoAre there any good quality OSS speech recognition libraries that are easy to get started with, or is it still so complex/expensive that this is a fantasy? I really love the idea of hacking something together for development so I don't need to use my arms and hands so much, or could lessen my mouse use (for accessibility reasons), but I don't know how realistic that is. Edit: I looked into cloud-based ASR, such as that provided by Azure and AWS, but that would mean network latency on top of the recognition latency, and that would drive me nuts!
- deleted 7y ago[deleted]
- wes-k 7y agoI dream of building a competitor to Siri and google and I’d probably use https://snips.ai/ https://snips.ai/. I think it gains recognition accuracy by having a limited skill set. Looks good though and has functionality for defining skills.
- jerf 7y agoSpeech recognition of a small set of possible words is basically a solved problem. It's why you've been able to call a phone support line and read numbers to it for years and years now; speech recognition for 10 digits and a handful of control words is basically done. So, if your project can be built on that, good news; you can build now.
- GordonS 7y agoThis looks interesting, as I don't need "conversation-level" ASR, I'd just need it to work with a limited grammar. But I'm unsure of what Snips actually is - I had a look at the website, but I don't know if this is OSS, commercial software, a library or what?
- ragebol 7y agoI've been trying out Snips, it's pretty cool and works reasonably well. A lot of the overall system is open source and runs offline, but the training happens on their servers and is closed source AFAIK. You download the trained model and can run it on a raspberry pi etc offline. But they claim that what Snips offers for free isn't nearly as good as what the commercial offering does. My priorities unfortunately shifted away from finding out about the quality difference. IIRC Snips's business model is to create custom voice agents that run offline for other companies and services around that, eg. custom hotwords etc.