3 ms·
If you use speech recognition systems on mobile (or a public web API), they are often handicapped due to space or processing constraints. Full on recognition mo
by kastnerkyle 11y ago
If you use speech recognition systems on mobile (or a public web API), they are often handicapped due to space or processing constraints. Full on recognition models in clean environments (no background noise, other speech, no compression e.g. something like a landline) backed by maximum hardware are quite good for English - we have much farther to go for other languages, which is one reason Baidu's work is so interesting.
Movie subtitles have poor alignment, usually contain multiple speakers (sometimes talking at the same time) and often contain sounds or other things which are not dialogue. Cleaning this is expensive, probably much more than just getting transcriptions of single speaker samples. It corresponds much closer to "real human life" but that is not where papers are published, unfortunately.
We have thousands of hours of ebooks (librivox - librispeech is a 1K hour subset used by many) that are used in open source speech recognition systems, and Baidu has a direct line on many more hours than that.
The better than human line is (in almost every paper - Baidu is not alone in this at all!) bullshit though - while the system is quite good, they only compare to humans given a fragment of a statement without context. Humans are very, very good at inferring given context (would you recognize speech the same way at a funeral/dinner party/college lecture/hospital?), and this is where most models fall flat.
Conditional language models can help this to an extent, but human adaptation and the ability to generalize will take a while to catch up with.
- est 11y ago> Cleaning this is expensive, probably much more than just getting transcriptions of single speaker samples So, deep learning to clearing background noises?