16 ms·
Speech-to-Text Benchmark: Framework for benchmarking speech-to-text engines
- malceore 8y agoThis basically seems like marketing paraded as research.
- barftransit 8y agoI've noticed I feel great when I fast for an hour every other day. Sometimes I'll supplement turmeric root.
- sctb 8y agoIf you won't post according to the guidelines we'll ban the account. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- barftransit 8y agotest to see if I'm banned
- gok 8y agoIf Mozilla’s DeepSpeech is getting a 30% WER on this test set with a >2GB model... something is very wrong.
- kenarsa 8y agoThis is a bit of surprise to me as well. That being said. The dataset is a tough one. I am ESL, but some of the accented examples I have a hard time understanding. Also the recordings are not near field and there is sometime some background noise.
- p1esk 8y agoWhat is the WER of your model on Librispeech?
- deleted 8y ago[deleted]
- glup 8y agoIn my experience DeepSpeech with the default Mozilla model + 5-gram is extremely sensitive to background noise. Fiddling with the hyperparameters, e.g. upweighting the language model, helps a little.
- kenarsa 8y agoThanks. Do you use the default parameters and tune from there?
- smt88 8y agoI'm surprised no one has questioned that this benchmark is published by the creators of Cheetah. The conflict of interest is extreme. Is anyone enough of a STT expert to weigh in? Why should I trust this?
- kenarsa 8y agoI agree with your point. That is why we open sourced the benchmark so you can verify it yourself. You just need an Ubuntu machine with python 3.6. I appreciate the question.
- smt88 8y agoIt's the methodology, not the results or the code, that I'm suspicious of. For a highly variable task like STT, I'm sure an expert could contrive a test that gives far better results for either of the other programs tested. That's why it would help to know why this test is comprehensive and representative enough to be considered unbiased or otherwise where its biases are. I don't have that expertise myself.
- p1esk 8y agoI think the biases are pretty obvious, but the most serious shortcoming of this benchmark is that their result (30% WER on CV) is not reproducible: it's not clear what they trained their model on, and the model itself is not available, so you just have to take their word for it.
- kenarsa 8y agoThanks for the comment. Just wanted to quickly clarify that the model is available here: https://github.com/Picovoice/stt-benchmark/tree/master/resources/cheetah https://github.com/Picovoice/stt-benchmark/tree/master/resou...
- p1esk 8y agoNo, "reproducible" means that I can train it on the same data you used, and get the claimed result. Anything other than that is taking your word for it.
- mhei 8y agoI wonder if they actually trained Cheetah on a different or the same dataset they are benchmarking on.
- kenarsa 8y agoThanks for comment. We don't train on the dataset being tested on. I am fairly certain the other engines don't as well.
- berbec 8y agoIf I'm reading this correctly, the resources used by Cheetah would allow your laptop to do stt for around 30 streams?
- kenarsa 8y agoIf you use the CPU that we used for testing yes. It's an Intel CPU. The detail is in readme. You can do 30 streams per core. If you have 2 cores you can do 60.
- berbec 8y agoWith such amazing efficiency, I would suggest forking off a second branch that massively increasing complexity for gains in your WER. CPU is just going to keep getting cheaper and faster, and being able to leverage the extra cycles for a platform that has them would allow you to dominate from embedded up to the Xeon space.
- kenarsa 8y agoThat is a good point. Definitely a valid roadmap and something we should consider. Thank you for the suggestion.
- btashton 8y agoI'm curious about the background of the team. They don't seem to have any linguistics experts from what I could tell.
- glup 8y agoIf that quote attributed to Fred Jelineck ("Every time I fire a linguist, the performance of the speech recognizer goes up") is right, they should have negative linguists.
- btashton 8y agoAnd if you blindly follow the machine learning path with no scientific context you will convince yourself that having negative people on a team is an answer that makes total sense. I worry about this in a lot of fields where we are heavily pushing machine learning. I was also not saying that this project is somehow poor quality, I just could not find much about the team and was curious about their research backgrounds.
- kenarsa 8y agoHello, sorry to interrupt the conversation. I am Alireza. I am the founder of Picovoice (maker of Cheetah). Thanks for mentioning the quote I heard it from few friends who used to work in Nuance. But didn't know where it came from originally. We have a very small engineering team who deeply understands machine learning and (embedded) software. I have invented/co-invented few US patents in speech processing/recognition prior to Picovoice. But have no academic background in linguistics. Having worked with computational linguists, they are definitely a solid plus to the team when their skill set is utilized correctly. I appreciate your question and curiosity. We hopefully will provide more information about the team on our website soon. Thank you.
- DougMerritt 8y ago> If that quote attributed to Fred Jelineck ("Every time I fire a linguist, the performance of the speech recognizer goes up") is right, they should have negative linguists. That's a really, really, really stupid quote that pushes the increasingly popular world view that experts only drag the world down on every subject. Linguists' contribution is unrelated to performance, positive or negative. Performance is the job of computer scientists designing architecture and of skilled programmers doing smart implementations of architecture. The performance would probably go up if the company fired all presidents/vice presidents/CFOs etc., too (giving implementors free reign, unconstrained by costs and schedules), but the company would probably go under. Every specialist has their own role to play. Thinking that subject domain expertise is a negative is indefensible.
- CommanderData 8y agoCan this be used to transcribe voice data in real time? I am building a docker image which will eventually accept in-browser audio via WebSockets outputting transcription in real time without needing Google WebSpeech. https://github.com/ashwan1/django-deepspeech-server https://github.com/ashwan1/django-deepspeech-server I planned to use DeepSpeech but this looks promising given it's low resources.
- themarkn 8y agoCould you tell me some more about this? I have used the web speech recognition API in Chrome (coming soon to Firefox also) to create a free project for real-time, editable, transcriptions in the browser that could be projected on a screen or subscribed to on a person's own device. I am frustrated by: - lack of cross-browser support on whatever device is generating the transcript. It will work from any Android phone but not from iOS - No offline functionality, because the work is actually not done locally - Poor accuracy.. we can correct this on the fly with the live editing, but it feels like it is not using the latest & greatest Google has to offer in terms of STT I don't know much about Docker or how to "use" a Docker image. Not asking you to teach me, but do you think your project would be useful for what I'm talking about? Currently using Firebase for the backend but it's really just exploratory at this point.
- iam-TJ 8y agoI too would like more information about both of your projects. I'm currently designing a digital technology platform for a charity for the blind in the UK and transcription of speech to text is one aspect I'm investigating, in addition to the more obvious text to speech.
- themarkn 8y agoHere's a video demo of what we are working on: https://youtu.be/xcUxd9sOkaM https://youtu.be/xcUxd9sOkaM There are a few features not listed (like exporting a correctly-formatted subtitle file) but you'll get the general idea. There's a link to an old demo in the video description.
- oulipo 8y agoHi, I'm a co-founder of https://snips.ai https://snips.ai and we are building a 100% on-device and private-by-design, open-source VoiceAI platform with our own ASR which works on Raspberry Pi3, Linux, iOS, and Android We already have a community of more than 14 000 developpers on the platform, you can get your own assistants running on a Raspbbery Pi in less than 1h, those are a few tutorials: - https://medium.com/snips-ai/voice-controlled-lights-with-a-raspberry-pi-and-snips-822e53d7ede6 https://medium.com/snips-ai/voice-controlled-lights-with-a-r... - https://medium.com/snips-ai/an-introduction-to-snips-nlu-the-open-source-library-behind-snips-embedded-voice-platform-b12b1a60a41a https://medium.com/snips-ai/an-introduction-to-snips-nlu-the...
- mcjiggerlog 8y agoI can't be the only one that sees "Token Sale" and immediately loses all interest.
- nathell 8y agoI wonder how these engines compare to cloud services described in [1]. [1]: https://blog.rebased.pl/2016/12/08/speech-recognition-1.html https://blog.rebased.pl/2016/12/08/speech-recognition-1.html
- tootie 8y agoThe cloud options almost certainly have better accuracy and use less memory but it's at the cost of latency, network dependency and usually price.
- glup 8y agoCommercial APIs from leading companies generally achieve better performance, but besides obvious price and network latency, they are complete black boxes so you can't diagnose and fix problems.
- kenarsa 8y agoGreat question. I believe that someone has already performed this measurement on a variety of could APIs (Google, Amazon. MS. etc.). I remember seeing it on GitHub while ago. I can't find it right now with a simple Google search. But I will spend more time later in the evening and comment here when I have it. The comments are absolutely correct. Could services (can) do better simply because they have access to more compute resources and also data (what is sent to them can be/is used for training later). There are situations where on-device is preferred due to privacy reasons, latency, cost, or lack of internet connection.
- DonHopkins 8y agoIt would be great to have a service that let you personally audition a bunch of speech recognition engines to measure and compare how well they understand YOUR voice.
- sonnyblarney 8y agoYeah, but it'd be like getting a blood test and trying to interpret it yourself :). There are so, so many variables - and many engines could be optimized for 'your voice and the things you're going to say right now given cpu/memory, quality of microphone, background noise etc..' The game of NLP is inherently about dealing with 'noisy channels' (in the academic sense) in which there is kind of a probabilistic guarantee of imperfection. So then it comes down to creating the best products in a given context, which is almost always 'less than optimized' for any individual. So there's the model size, cpu/ram, quality of signal (microphone, network) just to start. Optimizing for standard english means probably reducing quality for people with accents. Maybe in a specific context you could go from 80% accuracy to 85% accuracy for 'most of us ' - but then yo go from 60% to 40% for anyone with an accent. And if you reduce the accepted vocabulary, you can get way better results. Of course, we all might want to say words like 'obvolute' and 'abalienate' every so often. Kind of thing. It's really fun from an R&D perspective, but it's a product managers nightmare. Consumer expectations with these technologies are really challenging because of inherent ambiguity in a system, people kind of want perfection. And there are always corner cases where it would seem like things should be say, but they're not - because the word you're saying is common, and you'r saying it 'perfectly clear' ... but little do you know there are 2 or 3 other very rare words that sound 'just like that' ergo ... problems. From a product perspective, it basically always feels 'broken' which is such a terrible feeling :). But it can be fun if you like really hard product challenges which have less to do with tech and more to do with pure user experience, expectations, behaviours etc.
- glup 8y agoI think this is a great summary of where we are at now, but I think stronger broad-coverage language models (=expectations for what people say, better generative models of speakers) are feasible (for a couple 10s-100s of millions in R&D) to bring ASR up to parity with people. It's pretty clear we are getting close to the limits of what acoustics can offer, and it's the language model that is the next frontier both in terms of accuracy and real-time performance.
- tootie 8y agoWhat would you folks use for a much narrower use case more like an IVR system? Like if I could give a very restricted set of inputs to recognize and don't need much if any learning to happen.
- kenarsa 8y agoThanks for the question. You can use our first product: https://github.com/Picovoice/Porcupine https://github.com/Picovoice/Porcupine. It is a voice-control (wake-word) engine. It allows you to detect multiple keywords (phrases) in the audio stream in real time with no delay and it fully runs on-device. No cloud connection required.
- aviv 8y agoIBM Watson with https://freeswitch.org/confluence/plugins/servlet/mobile#content/view/16352039 https://freeswitch.org/confluence/plugins/servlet/mobile#con... Or stream to Google Cloud. Both are used in production successfully already. Latency is on par with what you would experience calling FedEx and talking to their ASR bot.
- el5r 8y agoLooks like Kaldi gets 4.27% WER on CV: https://github.com/kaldi-asr/kaldi/blob/master/egs/commonvoice/s5/RESULTS#L19 https://github.com/kaldi-asr/kaldi/blob/master/egs/commonvoi... But that's likely trained exclusively on CV audio and using a language model derived from the CV training data so that's also not a fair comparison unless the other engines were trained the same way. Comparing systems trained on different datasets (with different language models) like this is like comparing apples to oranges. Mixing wildly different CPU and memory requirements into the benchmarks just makes it worse. It should be fairly straightforward to do an unbiased comparison by training and evaluating with the standard Librispeech split and language model. It might be interesting to see how accuracy improves as the models scale up until they match the resource requirements of the other engines. That said, the speed and memory usage are impressive and I like the focus on very low resource environments. Seems like it has a lot of potential even if it may not be SOTA.
- kenarsa 8y agoThanks for the link. I'll be sure to look into it. It is impressive. A disclaimer is that we use the valid train portion of CV as part of our training set. But it is less than 10% of the train set (in terms of hours). Also, we do not employ an LM mainly because the systems we are targeting do not have enough storage for a strong LM (usually the storage on them maxes out at 64 MB). Cheetah is an end-to-end acoustic model. For later versions, we might be able to add a well-pruned LM for specific domains with limited vocabulary to boost the accuracy with limited storage available. I fully agree with your points. I am taking notes here as I think we should follow up on a couple of your suggestions. Scaling up to DeepSpeech model size can be a bit tricky as it would require much more compute resources (GPU). But should be quite doable with time and budget. Thanks again for your comments and suggestions. As you correctly pointed out our main focus is the very low resource (CPU/Memory) embedded systems.
- ausjke 8y agoJust checked out and tried to play with it, turns out this is a binary-only(closed source) release actually. I'm on a mips platform so there is no way I can test it there.
- kenarsa 8y agoUnfortunately, we only offer Ubuntu x86_64 at the moment. We are planning to add more platforms (Mac and Windows). I make a note of your request and add it to our todo list. Thank you.