4 ms·
This. The article quotes DeepSpeech, wavenets, LSTMs of all sorts; essentially, all the neural networks that scale terribly. DeepSpeech for example is pretty h
by priansh 7y ago
This.
The article quotes DeepSpeech, wavenets, LSTMs of all sorts; essentially, all the neural networks that scale terribly. DeepSpeech for example is pretty heavy and requires a decent GPU to get anywhere near a realtime factor of 1.
Meanwhile ASR through HMM's consistently hits realtime factors sub-1 and can run on small CPUs. e.g. the default models Kaldi ships with outperform DeepSpeech on a lot of modern examples, and are exponentially faster.
Moreover, training these HMM's is something that is feasible for a normal developer. Training newer models requires data of scale and quality (iirc Mozilla's models are trained on Common Speech which is an enormous crowd sourced dataset, and Google's wavenet models use an internal dataset of very high quality and quantity).
Until the models get more practically achievable, ASR for average people will probably continue to be dominated by Kaldi, Sphinx etc
- GordonS 7y agoDo you have any advice or links for a "normal" developer to get started with HMMs for speech recognition? I'd love to build something just for myself here.
- priansh 7y agohttps://github.com/jcsilva/docker-kaldi-gstreamer-server https://github.com/jcsilva/docker-kaldi-gstreamer-server ^^ If you're getting started, just following the steps here can get you set up really fast. IIRC there's also an HTTP endpoint at /recognize you can use instead of WebSocket if you're transcribing audio files so it's pretty cool!
- GordonS 7y agoNice, this doesn't look too bad to get started with - definitely going to give this a try as soon as I get a chance!
- lunixbochs 7y agoI'm hitting full-encode/decode realtime-factor in the ballpark of <0.01x on a quad core CPU with a tuned wav2letter
- Nimitz14 7y agoNote we're talking about LVCSR, you're not going to get 0.01 RTF doing that. Still, wav2letter is quite impressive from the numbers I've seen (although you're still going to need an insane amount of compute and data to train a good model). I've been meaning to try it out, but the setup is so complicated I haven't gotten around to it (I tried it and at some point while going down the dependency tree I said "fuck this" and stopped).
- lunixbochs 7y ago<0.01x in this case is fully encoding/decoding arbitrary input speech against a 700k word english wikipedia-based language model, for online/realtime use. I'd say it's pretty large vocabulary, and it's continuous in the sense that you can talk forever and it will output text regularly. (Wikipedia doesn't make the best LM, I just wanted to test with something that knew about a lot of interesting english words)
- Nimitz14 7y agoGoing to be honest but I don't think that is physically possible (i.e. I don't believe that, no offense). Unless you're using a very small beam and model.
- lunixbochs 7y agoOk, with a few caveats: 1. I don't have the quad core CPU we hit <0.01x on, I only have a dual core laptop CPU. (The quad core numbers were with an i7 7700k or so, they came from the other person I'm working on this with). 2. I'm not running a parallel decode. That's in a branch from my collaborator I haven't merged/built myself yet since I'm only on a dual core. 3. Screen capture has a CPU hit and seems to have slightly increased my RTF during recording. Here's your comment read with 0.05x-0.10x on a dual core CPU: https://youtu.be/jIgUKwR-LaA https://youtu.be/jIgUKwR-LaA Is that enough to convince you that with a stronger CPU and parallel decode we can hit 0.01x?