5 ms·
Does anyone know of other open-source projects in the speech-to-text space? DeepSpeech was one of the most promising projects, especially the latest versions...
by eruleman 6y ago
Does anyone know of other open-source projects in the speech-to-text space? DeepSpeech was one of the most promising projects, especially the latest versions...
- nshm 6y agoTry https://github.com/alphacep/vosk-api https://github.com/alphacep/vosk-api. It supports 10 languages, works on Android and RPi and also has big and more accurate server models. Other good ones are https://github.com/daanzu/kaldi-active-grammar https://github.com/daanzu/kaldi-active-grammar and https://talonvoice.com/ https://talonvoice.com/ There are toolkits for research like https://github.com/kaldi-asr/kaldi https://github.com/kaldi-asr/kaldi, https://github.com/espnet/espnet https://github.com/espnet/espnet, wav2letter, Espresso, Nvidia/Nemo, https://github.com/didi/athena https://github.com/didi/athena. You can try them too if you want to go deep. Some of them have interesting capabilities.
- posguy 6y agoComparing DeepSpeech v0.7.4 to Vosk using plain spoken English samples from male and female speakers, they seem to be performing the same if I use vosk-model-small-en-us-0.3 and the full size DeepSpeech model. When I use vosk-model-en-us-daanzu-20200328 the result is perfect on many of these tests, though it does not do punctuation or capitalization outside apostrophes. IIRC there is another project on Github that can add basic formatting though. I am quite surprised with vosk's performance, it even handles odd words like Puget Sound well! Need to test our more accented audio on it, but this is quite exciting.
- albertzeyer 6y agoThere are a lot of open source projects in this space. DeepSpeech is actually one of the outsiders (they are not represented well in the academic community), and also not quite competitive to other software (at least last time I checked). E.g. some very active projects are: * Kaldi (https://github.com/kaldi-asr/kaldi/ https://github.com/kaldi-asr/kaldi/) obviously, probably the most famous one, and most mature one. For standard hybrid NN-HMM models and also all their more recent lattice-free MMI (LF-MMI) models / training procedure. This is also heavily used in industry (not just research). * ESPnet (https://github.com/espnet/espnet https://github.com/espnet/espnet), for all kind of end-to-end models, like CTC, attention-based encoder-decoder (including Transformer), and transducer models. * Espresso (https://github.com/freewym/espresso https://github.com/freewym/espresso). * Google Lingvo (https://github.com/tensorflow/lingvo https://github.com/tensorflow/lingvo). This is the open source release of Googles internal ASR system, and used by Google in production (their internal version of it, which is not too much different). * NVIDIA OpenSeq2Seq (https://github.com/NVIDIA/OpenSeq2Seq https://github.com/NVIDIA/OpenSeq2Seq). * Facebook Fairseq (https://github.com/pytorch/fairseq https://github.com/pytorch/fairseq). Attention-based encoder-decoder models mostly. * Facebook wav2letter (https://github.com/facebookresearch/wav2letter https://github.com/facebookresearch/wav2letter). ASG model/training. * (RETURNN (https://github.com/rwth-i6/returnn https://github.com/rwth-i6/returnn) and RASR (https://github.com/rwth-i6/rasr https://github.com/rwth-i6/rasr), our own, although this is currently free for academic use only. It is used in production as well. Supports hybrid NN-HMM, CTC, end-to-end attention-based encoder-decoder, transducer, etc.) And there are much more. You will also find lots of ready-to-use trained models.
- Bootwizard 6y agoCan you run audio files through any of these or do they only support audio from microphones?
- nmstoker 6y agoAt the point of them taking in input to process, audio that comes from a microphone or comes from a file is basically just a series of numbers and is the same. So there's no barrier in terms of feasibility. Whether they're all set up to do that "off the shelf" is a different matter but it should be fairly straightforward to add this to any that lack it and because they're open-source anyone could do a bit of Googling etc and find suitable code to adapt to do it. I know DeepSpeech definitely can take audio from files directly as input as I've used it that way before, and I strongly expect many (or possibly all) of the others could too.
- posguy 6y agoDeepSpeech and Vosk can accept audio files, although each wants them formatted in a slightly different mono WAV format. See my other comment for a comparison of the two: https://news.ycombinator.com/item?id=24248238 https://news.ycombinator.com/item?id=24248238
- convery 6y agoYou seem to know a lot about the topic, any idea about the current state of text-to-speech? Haven't seen any opensource projects that would make, for example, an ebook enjoyable.
- kouohhashi 6y agodeepspeech.pytorch is a good one. Since Mozilla's DeepSpeech project is still using tensorflow 1.x, I think pytorch implementation is actually better. https://github.com/SeanNaren/deepspeech.pytorch https://github.com/SeanNaren/deepspeech.pytorch