3 ms·
Hi, my name is Alexander, I one of authors of both Gradient pieces, Open STT and silero.ai. We are planning on adding 2-3 new languages soon (English, German,
by snakers41 7y ago
Hi, my name is Alexander, I one of authors of both Gradient pieces, Open STT and silero.ai.
We are planning on adding 2-3 new languages soon (English, German, maybe Spanish), so if you would like us to fully open source everything - please support us here - https://opencollective.com/open_stt https://opencollective.com/open_stt
A lot of feedback here, very nice!
As with the Russian Speech community for some reason people are willing to share their feedback on third-party forums, but do not bother to write to the author directly.
If you have something to say - please email me - aveysov@gmail.com or telegram me (@snakers4 or @snakers41) if you would a faster reply.
Also it seemed to me that some of the readers missed that there were actually 2 parts of the article - https://thegradient.pub/towards-an-imagenet-moment-for-speech-to-text/ https://thegradient.pub/towards-an-imagenet-moment-for-speec....
With this all addressed, to answer some common criticisms / points raised below:
> Kaldi cool, NNs not cool
Though it is true that production solutions CAN be built with Kaldi (and I have even spoken with people who built them, as well as vosk author), we purposely omitted Kaldi because we believe that it is a technological dead-end. Also it depends a lot on antiquated technology, is very difficult for new-comers etc etc
Another problem is that ... it requires a lot of specific knowledge to add languages. Whereas our approach is just plug-and-play - just add more data!
For example - our best model in production is just 300-500 lines of code in PyTorch.
With so much resources poured in tools like PyTorch, it just makes sense to use tools that are simple / robust / offer a lot of future-proofing and cool features etc etc Also you can ofc switch DL frameworks if you wish =)
Also you can do CV, NLP, etc with PyTorch unlike Kaldi.
As someone noted, does not really matter what you are using, if it suits your compute.
Fully e2e approaches still are GAFA scale only, and I just realized why Google even bothers with them - please see my post here https://t.me/snakers4/2445 https://t.me/snakers4/2445
> ML APIs changing w/o backwards compatibility
TF does that, and it is a joke. Paid marketing says that TF is cool, but real practicioners (that I know personally) use PyTorch.
It is a holywar, but I believe TF just has a lot of captive audience ...
PyTorch core API on the other hand has been mostly stable since 0.4 (now it is on 1.4)! (!!!) No one speaks about this for some reason.
> Deep Speech is a bad starting point
Yes and no. Vanilla huge LSTMs are really GAFA scale only.
But out networks is at least an order of magnitude faster.
Please see this - https://thegradient.pub/towards-an-imagenet-moment-for-speech-to-text/ https://thegradient.pub/towards-an-imagenet-moment-for-speec...
> cool open source datasets
Check out Open STT!
We have not been able to OSS everything, because Russian corporations would use it w/o even crediting us.
But you can support us and change it. Also other languages.
caito.de is really cool, albeit very small
common voice - is very call, but on the smaller side
libri-light - you have to align it yourself, so no difference really - easier to make it from scratch
Also I do not understand why people use flac, when opus exists?
- woodson 7y agoHi Alexander, I appreciate you posting here; I didn't feel too strongly about this to sign up for an account to comment directly on TheGradient. > Kaldi cool, NNs not cool It seems you're misrepresenting things. Kaldi models are NNs, they are trained using a sequence-level objective function (LF-MMI, not entirely unlike CTC), they can also be trained completely end2end (without using HMM-GMM systems to generate alignments for CE regularization and obtaining lattices). The Kaldi authors are currently working on moving NN training to PyTorch. The fact is that, if you need a complete easy-to-use off-the-shelf solution ("Just give me the transcription of that audio file!"), Kaldi is not it. You need to train models and you need to use other projects, such as vosk, kaldi-gstreamer-server, etc. or build your own solution. People have done so and use it in production successfully. Use the right tool for the job, a poor workman blames his tools, yadda yadda. Perhaps it's just that speech recognition IS difficult, there is no one-size-fits-all-use-cases solution, and when people can't figure out Kaldi, they seem to think they need to throw out everything and pray that CTC will fix it somehow. Get sufficient data in your domain and things might work better. Don't use a couple of audiobooks for training, expecting that it will perfectly transcribe Indian accented English callcenter conversations. I really appreciate your effort on releasing OpenSTT, in particular that you provide data in different conditions and domains (phone calls, lectures, broadcast radio, etc.). This is useful no matter what particular tool (Kaldi, DeepSpeech, w2l, RNN-T, transformers) people are using.