5 ms·
Hey HN! I've been working on this for a while now. While there are other on-premise solutions using older models such as DeepSpeech [0], I haven't found a depl
by iceychris 6y ago
Hey HN!
I've been working on this for a while now.
While there are other on-premise solutions using older models such as DeepSpeech [0], I haven't found a
deployable project supporting multiple languages using the recent RNN-T Architecture [1].
Please note that this does not achieve SotA performance.
Also, I've only trained it on one GPU so there might be room for improvement.
Edit: Don't expect good performance :D this is still in early stage development. I am looking for contributers :)
[0] https://github.com/mozilla/DeepSpeech https://github.com/mozilla/DeepSpeech
[1] https://arxiv.org/abs/1811.06621 https://arxiv.org/abs/1811.06621
- the_biot 6y agoWhy on earth would you put a demo front and center that shows your software doing a terrible, terrible job?
- brmgb 6y agoThat's called setting up expectations. If you know your project might interest people but needs work, why pretends it's good when it's not? They seem to be courting contributors more than users anyway. I found the video to be funny. It nicely highlights both the current limitations and the ambition of the projects. Bold choice certainly but I think it works.
- dcsan 6y agoand it's maybe a dig at Macron's accent at the same time :D although the author is a student in germany. Anyway you should join the Discord, we discussed this there too... https://discord.gg/pqTMeP5D3g https://discord.gg/pqTMeP5D3g
- th3h4mm3r 6y agoHi! What should you need to implement other language i.e. Italian or French? I mean: it's a problem due to the less of datas or what? Another question: could you use for example mozilla voice data to train/test?
- iceychris 6y agoData and compute are the largest hurdles. I only have one GPU and training one model takes 3+ days, so I am limited by that. Also, scraping from YouTube takes time and a lot of storage (multiple TBs). Mozilla Common Voice data is already used for training.
- jack_pp 6y agoWhy does it take a lot of data? Afaik you can select lower quality in youtube-dl but you don't even need video do you?
- klysm 6y agoI know you can scrape only audio from YouTube with YouTubeDL but it’s somewhat annoying
- Shared404 6y agoI use something akin to 'alias downloadmusic='youtube-dl --extract-audio --audio-quality 0 --extract-metadata' in my .bashrc I find that helps with the annoyance of downloading things off of YT. This is for music obviously, but there's an option to download subtitles as well. EDIT: Typed this from memory, there may be errors in the alias.
- jerf 6y agoyoutube-dl -f bestaudio $URL Dunno when that went in but it works now.
- whimsicalism 6y ago> Why does it take a lot of data? Afaik you can select lower quality in youtube-dl but you don't even need video do you? But you need supervised data too.
- th3h4mm3r 6y agoSo do you scrap videos from youtube with subtitles to collect data?
- whimsicalism 6y agoAwesome project - I'm also working on a similar idea for an on-premise ASR server! Any reason you decided to go with RNN-T?
- woodson 6y agoYou can also check out https://github.com/TensorSpeech/TensorFlowASR https://github.com/TensorSpeech/TensorFlowASR for inspiration (not my project, not involved). It implements streaming transformers and conformer RNN-T (but in TF2). Deployment on device as TFLite. So far, there aren't many usable pretrained models available (just LibriSpeech), but with some work it could turn out quite nicely.