3 ms·
Dylan from AssemblyAI. We're working on this! We didn't plan to launch on HN tonight which is why the benchmarks aren't ready, but we know we definitely need th
by dylanbfox 8y ago
Dylan from AssemblyAI. We're working on this! We didn't plan to launch on HN tonight which is why the benchmarks aren't ready, but we know we definitely need these for the community.
We've found most public benchmarks, especially Libri, are not that representative of real world data we see in production. Most real world data we see is a lot noisier, and has worse recording quality like low bitrates and compression from mp3 encoding.
We do worse than state of the art benchmarks on Libri Clean today, for example (I think we are around 7% WER last time I checked), but are much more accurate on real world data than models reporting 3-5% WER on Libri. This is why we want to make sure we are thorough when we report our benchmarks on popular datasets like Libri.
- braindead_in 8y agoYeah, I would agree that the academic datasets are not representative, especially for STT. Our models do around 8.7% on LibrisSpeech Clean. But on our internal dataset PaddlePaddle's WER is 29% whereas we do around 15%. And we regularly see higher WER's in production, especially for accented and noisy files. Hopefully continuous re-training will help improve the generalization. Here's the output from our model the 6063 file. 0:00:00.7 S1: I'd say there's a such thing as eating too much, but I just have a massively fast the table of them and so, I constantly eating so. Do you do diarisation and punctuations as well?
- dylanbfox 8y agoRight, we noticed similar findings. We do automatic punctuation now, and do diarization when there is more than one channel in the audio file. We're launching diarization on single channel audio with multiple speakers very soon. We're currently focused on improving some of our customization features, and then we plan to ship single channel diarization. Thanks for sharing your results! We have more samples here if you want to do more comparisons: https://blog.assemblyai.com/2018/08/09/cutting-edge-phone-call-transcription-with-assemblyai/ https://blog.assemblyai.com/2018/08/09/cutting-edge-phone-ca...
- braindead_in 8y agoThat's great. We have really struggled with diarization. None of the systems out there actually work! We get close but still mess it up from time to time. Here are the results on the other files. 4333.mp3: Oh yeah, it's still pretty tight, though. It's very challenging. They actually pull everything out of your. 7510.mp3: Demons on TV like that. And for people to expose themselves to being rejected on TV or humiliated by fear factor or. 8036.mp3: Well, I feel like as far as... as far as cursing and language, because I feel like as long as it's not necessarily in context, but. 8522.mp3: Stuff to you, so you don't have to spend any body. He.
- dylanbfox 8y agoThanks for sharing this! It's awesome to see another independent company tackling STT well. If you guys ever want to try to collaborate, let me know! My email is in my profile.
- sushanthiray 8y agoI've been working on the problem of speaker diarization in the wild for better part of this year. Would love to have a chat with you to see if our diarization utility can help you guys out. Here is the link: https://www.deepaffects.com/apis https://www.deepaffects.com/apis My email is in my profile.
- woodson 8y agoHow about the HUB5 eval2000 set or RT-03 for conversational telephone speech?
- dylanbfox 8y agoThose test sets are definitely more real-world than Libri Clean in our experience, but still are not as "dirty" (noise, muffling, cross talk, recording quality, etc.) as we see in production.
- woodson 8y agoTrue, it's not nearly as dirty as real-world phone calls. Even when cross talk is not an issue, e.g. in voice mails, results on eval2000 are not representative of the performance.