4 ms·
Some languages are significantly easier to train against than English, for example 80hrs of Turkish got some decent results for this group using Mozilla Deepspe
by StudentStuff 8y ago
Some languages are significantly easier to train against than English, for example 80hrs of Turkish got some decent results for this group using Mozilla Deepspeech: https://arxiv.org/pdf/1807.00868.pdf https://arxiv.org/pdf/1807.00868.pdf
For context, Baidu used a bit over 5000 hours of English to get a decent model.
Also, OP may have a pronunciation difference (regional, hyper-local or just from learning English as a second language) that isn't handled well by Google's models.
- conistonwater 8y agoI suspect deep learning researchers have no real incentive to use the minimum amount of data sufficient to get their thing working. The worst case for them is if their model underperforms because of insufficient data instead of something truly DL-related, so they use as much data as they have/can handle. It could still be that you're right, but I wouldn't rely on those kinds of published numbers.