3 ms·
I don't get the claim about lack of training data. Most TV stations broadcast with closed captioning and most movies have subtitles available, which should giv
by devit 6y ago
I don't get the claim about lack of training data.
Most TV stations broadcast with closed captioning and most movies have subtitles available, which should give millions of hours of training data.
It's technically copyrighted, but as long as you don't distribute the video as well, they aren't going to care about it.
Also the complaint that the data lacks compression artifacts is completely ridiculous and absurd: if you want compression artifacts, just compress and decompress the speech yourself!
You can also pay people to transcribe audio or read text (or even do it personally), and since this is something anyone can do it can be paid at very low rates.
- PeterisP 6y agoThere's lots of resources for English and much less for many other languages - this article is about the non-English case. Closed captioning and subtitles are often used, but they are 'dirty data' - they are usually not a one-to-one match, the differences mess up training. And licensing issues are a pain; "they aren't going to care about it" is not a solution and things like background music in movies (the rightholders very much do care about it, even if just to make a point) make it pretty much impossible for e.g. some university to legally distribute a dataset that includes movie audio tracks; so they can run some experiments on it themselves but as soon as you want any collaboration, that data is taboo. Paying people to transcribe or read works, but it's not cheap. Reading 10000 hours takes at least 10000 hours of paid work and generally a bit more than that - mistakes matter, so you need review and correction. If you try transcribing things yourself, a tiny 100 hour dataset is going to take you at least something like 400 hours, which is months of work. So that's the point - getting a usable dataset for some language costs hundreds of thousands of dollars if you're frugal, and millions if you want good results. And there are very many languages in the world.
- eindiran 6y agoPetrisP's answer covered a lot of your question, but there are some additional considerations, even for English. Let's say you decide to grab all of those TV broadcasts and use them as training data (we'll ignore the fact that very little speech is as clean as the speech in TV broadcasts). Every mistake that occurs in the transcriptions you're using for training data represents the potential for mistraining your model when you perform a forced alignment [0]. Since most forced alignment tools perform a feature transformation step to increase how discriminative the phones are in feature space (ie how different they are from each other) and a step to reduce the number of dimensions, small errors in the transcriptions can result in huge differences in decoder behavior. Now factor in that your lexicon (the set of pronunciations you're using to map the orthographic transcriptions onto the possible phoneme strings that the speaker could have actually said) is almost certainly incomplete and very likely doesn't include nearly enough pronunciations for speakers with non-standard dialects or speakers with accents. Plus there are all sorts of typos in the original transcriptions; this often means that you need to run g2p to generate shitty pronunciations for words that actually don't exist, adding further errors into your model. Suddenly, you start thinking that all these decoder problems you're seeing are problems with the training data as it currently exists and the only solution is to do a phonetic transcription of the training data! You can still get that done, but since you've realized some of the blame lies with your lexicon you can't just dump the videos onto Mechanical Turk and have random people do them: you now need to get people that are familiar with phonetic transcription. And you decide that since even experts mess up a fair amount when transcribing speech, you're going to get 2 people to transcribe everything and throw away everything where they disagree. You aren't Google, so your training budget isn't that big. For each hour of speech you're spending 2 expert man hours and still throwing away 6 minutes and you need 10K+ hours of data at the bare minimum! And looking at the literature, it looks to you like there is a near linear correlation between who has the most data and who has the models that perform the best! So that's why people end up claiming that a fundamental problem with building out ASR is a lack of training data: because getting it and making it usable is very expensive, and because having lots of good training data let's you avoid addressing other difficult issues in ASR. [0] http://www1.icsi.berkeley.edu/Speech/faq/forcedalign.html http://www1.icsi.berkeley.edu/Speech/faq/forcedalign.html