6 ms·
Going along with this: What are the latest and greatest open source speech-to-text models and/or tools out there? Would love to hear from experienced practitio
by jalopy 5y ago
Going along with this: What are the latest and greatest open source speech-to-text models and/or tools out there?
Would love to hear from experienced practitioners and a bit of detail on the experience.
Thanks HN community!
- orra 5y agoMozilla announced Deep Speech[1] around the same time as Common Voice. Mozilla Deep Speech is an open source speech recognition engine, based upon Baidu's Deep Speech research paper[2]. Unsurprisingly, Deep Speech requires a corpus such as... Common Voice. [1] https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech [2] https://arxiv.org/abs/1412.5567 https://arxiv.org/abs/1412.5567
- zerop 5y agoVosk is my favourite. I have used deep speech too. Vosk works better.
- nshm 5y agoThank you. I deeply appreciate you mention our efforts. We spend quite some time and knowledge to build accurate speech recognition. Not that easy to get as much mentions as Mozilla, so we are thankful for every single one!
- zerop 5y agoVosk just works good and it works on mobile platforms too. One suggestion is to put lisence on alphacephei site. GitHub repo has it, but site doesn't.
- thom 5y agoSame question for text-to-speech!
- kcorbitt 5y agoI've had good results with https://github.com/flashlight/flashlight/blob/master/flashlight/app/asr/README.md https://github.com/flashlight/flashlight/blob/master/flashli.... Seems to work well with spoken english in a variety of accents. Biggest limitation is that the architecture they have pretrained models for doesn't really work well with clips longer than ~15 seconds, so you have to segment your input files.
- mazoza 5y agohttps://github.com/coqui-ai/STT https://github.com/coqui-ai/STT
- jononor 5y agoHave used VOSK a bit recently. The out-of-the-box experience was great compared to earlier projects (looking at you Kaldi and Sphinx...). Word-level audio segmentation was one usecase, https://stackoverflow.com/a/65370463/1967571 https://stackoverflow.com/a/65370463/1967571
- woodson 5y agoNVidia NeMo: https://github.com/NVIDIA/NeMo https://github.com/NVIDIA/NeMo
- blackcat201 5y agoI created edgedict [0] a year ago part of my side projects. At that time this is the only open source STT with streaming capabilities. If anyone is interested the pretrained weights for english and chinese is available. [0] https://github.com/theblackcat102/edgedict https://github.com/theblackcat102/edgedict
- lazyresearcher 5y agoKaldi and DeepSpeech both support streaming, right?