Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
iceychris
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
iceychris
5y ago
I'm using NixOS with i3 as my daily driver, can recommend.
2.
▲
by
iceychris
6y ago
Depending on your objective, noisy data might be useful. I'd like LibreASR to also work in noisy environments, so training on data that is noisy should already help a bit with that. But yeah - stammering and delays are present not only
3.
▲
by
iceychris
6y ago
Audio boards based on ESP32 boards are quite under the radar and have lovely features for just a few bucks. Running LibreASR on a RPi should also be feasible soon. Thank you for your kind words! :)
4.
▲
by
iceychris
6y ago
Data and compute are the largest hurdles. I only have one GPU and training one model takes 3+ days, so I am limited by that. Also, scraping from YouTube takes time and a lot of storage (multiple TBs). Mozilla Common Voice data is already us
5.
▲
by
iceychris
6y ago
I haven't trained on LibriSpeech exclusively, but yes, the perf on LibriSpeech dev is quite bad, around ~60.0 WER. If the poor alignment of yt captions is the issue, maybe concatenating multiple samples helps a bit.
6.
▲
by
iceychris
6y ago
Yes, probably. The data I trained on mostly reflects UK and US accents.
7.
▲
by
iceychris
6y ago
Hey blackcat! Your project [0] helped me a lot! Pre-training the encoder sounds great, I'll maybe add it in the future. [0] https://github.com/theblackcat102/Online-Speech-Recognition
8.
▲
by
iceychris
6y ago
I have not yet trained a french model. Also, the gif shows Macron speaking to the congress with his english accent [0] [0] https://www.youtube.com/watch?v=RqUc1h7bZQ4
9.
▲
by
iceychris
6y ago
Right, fixed it, thank you :D
10.
▲
by
iceychris
6y ago
As I commented above, very poorly. It's still early days.
11.
▲
by
iceychris
6y ago
LibriSpeech, Tatoeba, Common Voice and scraped YouTube videos.
12.
▲
by
iceychris
6y ago
The upper transcript is YouTube's automatic transcription. Below is the web app transcribing live. And yes, it is actually missing a few words.
13.
▲
by
iceychris
6y ago
Hey HN! I've been working on this for a while now. While there are other on-premise solutions using older models such as DeepSpeech [0], I haven't found a deployable project supporting multiple languages using the recent RNN-T Arc
14.
▲
Show HN: LibreASR – An On-Premises, Streaming Speech Recognition System
(github.com)
233 points
by
iceychris
6y ago
|
71 comments
15.
▲
by
iceychris
8y ago
username: iceychris Thank you very much :)