3 ms·
One of the main factors for this is probably due to dataset size. Commercial STT models are trained on 10s of thousands of hours of data from real-world data. E
by dylanbfox 6y ago
One of the main factors for this is probably due to dataset size. Commercial STT models are trained on 10s of thousands of hours of data from real-world data. Even a decent model architecture is going to perform pretty well on that much data.
Most open source models are trained on Libri, SWB, etc. which are not really big or diverse enough for real-world scenarios.
But to max-out results the devil is in the details IMO (network architecture, optimizer, weight initialization, regularization, data augmentation, hyperparam tuning, etc) which requires a lot of experiments.
- bginsburg 6y agoThere are new, very large public English speech datasets: Mozilla Common Voice, National Speech Corpus, which can be combined with LibriSpeech to train large models.
- eindiran 6y agoIf you combine them you get 5ish K hours of speech for English, which is still fairly small compared to what most big players have access to.
- solidasparagus 6y agoAmazon worked with 7k hours of labeled data + 1 million hours of unlabeled data - https://arxiv.org/pdf/1904.01624.pdf https://arxiv.org/pdf/1904.01624.pdf