3 ms·
It's great to see innovation in the space of open data. There have recently been a number of assertions that better quality ML data will outperform better ML a
by 16bytes 8y ago
It's great to see innovation in the space of open data.
There have recently been a number of assertions that better quality ML data will outperform better ML algorithms, and this has certainly been true in my experience as well, especially in domains like speech recognition.
There's going to be a long road to catch up to the big players, however. Even 15 years ago there were companies who were doing 1M minutes of labled voice data per year.
The data gap between established players and newcomers to the market will continue to grow unless we invest in efforts like this.
- novaRom 8y agoOriginal Deep Speech 2 paper released few years ago mentioned a hundred thousand or two hundred thousand hours, but the amount of data has increased since then significantly. Still quite a lot of languages have very tiny datasets of transcribed data.