6 ms·
What I would like to see you would be an open source speech recognition system that is easy to use, and works. There is really no match for the proprietary solu
by johnramsden 9y ago
What I would like to see you would be an open source speech recognition system that is easy to use, and works. There is really no match for the proprietary solutions in the open source world right now, and that is disappointing.
Not to say that there aren't open source speech recognition systems, they just aren't completely usable in the way proprietary solutions are. A lot of research goes into open source speech recognition, where they are lacking is in the datasets and user experience.
Hopefully, Mozilla's Common Voice project https://voice.mozilla.org/ https://voice.mozilla.org/ will be successful in producing an open dataset that everyone has access to and will spur on innovation.
- LandoCalrissian 9y agoI recently have been trying out open source solutions for voice recognition for a personal project and you are very correct that it lags very far behind proprietary solutions. Pocketsphinx is still very limited and Kaldi takes quite a bit to setup in a usable fashion. There were a few other options I looked at that I can't think of from the top of my head but were all in similar condition. As the article says latency is still a problem and it's a huge problem in current open source solutions, some stuff I was testing was easily 5 seconds. I know that can be improved with configuration, but when dealing with libraries of 10 words or so, that's pretty bad. I feel like anyone who is seriously interested in this space has been scooped up by all the big companies and the open source solutions have really seemed to linger because of it. It's one of the first areas I've seen where open source alternatives are really behind the proprietary solutions. Kind of bummed me out.
- dharma1 9y agoKaldi is the best, there was just Tensorflow integration added which will hopefully speed up development (though I haven't seen any pretrained models for that yet). Here's a blog post - http://www.googblogs.com/kaldi-now-offers-tensorflow-integration/ http://www.googblogs.com/kaldi-now-offers-tensorflow-integra... The easiest way to deploy Kaldi is this - https://github.com/alumae/kaldi-gstreamer-server https://github.com/alumae/kaldi-gstreamer-server (or a docker image of the that)
- adrianbg 9y agoUnfortunately that Tensorflow integration didn't include acoustic modeling, so you still need to use Kaldi's neural net toolkit for that.
- LandoCalrissian 9y agoThat's the option I was looking at using, so glad to see I'm on the right track. Thanks!
- ddevault 9y agoI agree - I don't care how good Google gets it, this is an unsolved problem until I can do it with open source tools operating without an internet connection.
- dragontamer 9y agoGoogle has basically invented a special processor with a very, very, VERY weird architecture for these sorts of tasks: https://drive.google.com/file/d/0Bx4hafXDDq2EMzRNcy1vSUxtcEk/view https://drive.google.com/file/d/0Bx4hafXDDq2EMzRNcy1vSUxtcEk... I don't think this level of computational power can be achieved on a modern CPU, or even a GPU! But GPUs are probably the closest analog to Google's absurdly parallel architecture. To get a GPU working at maximum performance, you either have to go OpenCL2.0 or CUDA. Compared to OpenCL1.2, OpenCL 2.0 has a better atomics model, dynamic parallelism (kernels that can launch kernels), shared memory, and tons of other features. NVidia of course supports those features in CUDA, but NVidia's OpenCL support is stuck at 1.2. So in effect, CUDA and OpenCL are in competition with each other. Anyway, that's the current layout of the hardware that's available to consumers. I think its reasonable to expect a graphics card in a modern machine, even Intel's weak integrated-GPUs have a parallel-computing advantage over a CPU. So for high-parallelism tasks like audio analysis or image analysis, it only makes sense to target GPUs today.
- gok 9y agoThe TPU architecture isn't that weird...it's basically a hardware implementation of matrix multiplication. It also isn't a silver bullet for ASR, where neural networks are usually only used for a part of the recognition process.
- IshKebab 9y agoNeural networks are used for nearly all of ASR now. Last I heard only the spectral components were still calculated not using a neural net and the text-to-speech is now entirely neural network (i.e. you feed text in and get audio samples out). I'd be surprised if they don't do that for ASR too soon if they haven't already.
- j_s 9y agoI too would be interested in pointers to the leading open source options. Just yesterday there was a Show HN built with the https://github.com/kaldi-asr/kaldi https://github.com/kaldi-asr/kaldi project, emscripten-ized: https://news.ycombinator.com/item?id=15534531 https://news.ycombinator.com/item?id=15534531
- ghaff 9y agoWhen I looked a while back, CMUSphinx seemed to be the most promising option but I struggled to get it installed and got distracted with real work. Some discussions online suggest it’s still fairly poor compared to the online engines. Snips was mentioned here recently but I haven’t taken a look at it.
- CaptSpify 9y agoI had a working cmusphinx setup at one point. It was so bad that I eventually just tore it down. It showed promise, but I don't know if any work is being done on it.
- adrianbg 9y agoCMUSphinx is really old. Kaldi is hard to use, but it's much better.
- nmcfarl 9y agoThe first time I read this I misunderstood it's meaning. I now believe the parent is discussing the differences in usability not output. In which case I completely agree. However I will say that for my company's use case a properly configured sphinx install produces better results than kaldi. However, getting to a point where you can say that was not an easy task. Additionally, I actually believe that for most workloads that kaldi is likely better. Not ours though.
- dsacco 9y agoIs it an issue of open source software being inadequate, or is it the lack of sufficient training data and compute power local to your home? I’m under the impression that Google is mostly dogfooding its open source tooling for machine learning in GCP, and actually differentiates based on trained models and compute power.
- backpropaganda 9y agoA bit of both. When data doesn't exist, people aren't motivated enough to create the open source tools which would leverage it. While Librispeech is okay for academic research, it's not enough to create a good production-grade speech recognition system.
- adrianbg 9y agoThe problem wrt open source / free solutions is data. Kaldi is open source and gets state of the art results -- but the data costs a lot of money. Training the models is doable on a commodity GPU although it takes quite a while.
- jjwiseman 9y agoI didn't realize Kaldi could get state of the art results. Do you say that because you know of people doing that, or is your comment based on knowing the architecture of Kaldi?
- adrianbg 9y agoThe 8.5% in this file is what you'd compare to Microsoft and IBM's recent ~5% results. https://github.com/kaldi-asr/kaldi/blob/master/egs/fisher_swbd/s5/RESULTS#L103 https://github.com/kaldi-asr/kaldi/blob/master/egs/fisher_sw... Kaldi hasn't been in first place on that dataset recently, but it was a few years ago. On other more researchy datasets (eg. for distant speakers or languages other than English), the best system is often based on Kaldi.
- woodson 9y ago
- drzaiusapelord 9y agoMozilla may not be able to make much headway until these patents expire: https://www.quora.com/Is-anyone-working-on-an-open-source-version-of-Siri?share=1 https://www.quora.com/Is-anyone-working-on-an-open-source-ve... Apparently, the business model for voice services in the past has been to snap up as many broad patents as possible to keep competitors at bay. I read an interview with a google engineer a couple years back claiming the same thing. They have to carefully work around a patent minefield with their own services and how these patents are holding back better voice search on mobile technology.
- mintplant 9y agoSpeech synthesis is in a similar situation. The FOSS options that I know of are completely primitive compared to their proprietary counterparts, particularly those locked away behind the cloud. Unfortunately this area seems to receive much less attention than speech recognition (AFAIK it's a non-goal of the Mozilla project, for example).
- ghaff 9y agoI suspect that it’s widely recognized that getting incremental advances in speech synthesis is really hard and, unlike potentially speech recognition, it doesn’t really solve a problem that people have. For business services and consumer devices slightly better voices are valuable but, for example, the improvements in the latest Siri voice don’t make any new things possible. But it’s part of the fit and finish of an expensive phone.
- deleted 9y ago[deleted]
- adrianbg 9y agoYou are correct that the problem is data. Kaldi is hard to use, but making it easier to use isn't as hard as getting good training data. Mozilla's project is a good start for some purposes. One flaw with it is that they're having people read sentences. When people read, they tend to speak more clearly than when they're figuring out what to say on the fly. This means models trained with Mozilla's data will tend to need people to speak extra clearly than eg. a model trained on conversational data. (edits: spelling/grammar)
- pishpash 9y agoOh it's not "hard" to get training data, you just need loads of money to buy the existing datasets.
- adrianbg 9y agoHah sure. By "hard," I meant that it's the largest hurdle. And probably even the common research datasets aren't enough to give you results competitive with Google, etc. AFAIK Google uses its own hand-transcribed data.
- dharma1 9y agoor have more effective ways to collect tons of open source speech data. The Mozilla Common Voice project is really cool, but they should make it way easier for people to contribute. Like, adding a mic button for voice search next to their main search toolbar on Firefox, and then ask for permission to use that data for research.
- flukus 9y agoIt's only a data issue if they're trying to reproduce google voice. Personally I got much better results from dragon naturally speaking 20 years ago than I get from google voice today, the cost was that you had to train it yourself first, but the benefit was it was trained for you, not "everyone". The later is the approach I'd prefer to see mozilla/OSS take.
- adrianbg 9y ago
- oulipo 9y agoI'm the cofounder of Snips.ai and we are building a 100% on-device Voice AI platform, which we want to open-source over time You can build your voice assistants and run them for free on a Raspberry Pi 3, or Android
- allenleein 9y agoOpen source speech recognition on Github: 1. speech-to-text-wavenet (https://github.com/buriburisuri/speech-to-text-wavenet https://github.com/buriburisuri/speech-to-text-wavenet) 2. kaldi(https://github.com/kaldi-asr/kaldi https://github.com/kaldi-asr/kaldi) 3. Speech recognition module for Python (https://github.com/Uberi/speech_recognition https://github.com/Uberi/speech_recognition) 4. DeepSpeech(https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech) 5. Natural Language Processing Tasks and References(https://github.com/Kyubyong/nlp_tasks https://github.com/Kyubyong/nlp_tasks) FYI.