6 ms·
I'm about to start as a professor in CS education, and am hoping we're getting close to the point where I can easily transcribe interviews and high-quality dial
by cproctor 7y ago
I'm about to start as a professor in CS education, and am hoping we're getting close to the point where I can easily transcribe interviews and high-quality dialogue audio using open-sourced models running on machines in my lab. I'm tired of paying $1/minute for human transcription that's not great anyway, and would love to undertake research that would require processing a lot more audio than is affordable on those terms.
I haven't kept up with developments over the last two years--anyone have a sense of whether this is close to being a reality?
(I've taken a bunch of Stanford's graduate AI courses on NLP and speech recognition; I can read documentation and deploy/configure models but don't have much appetite for getting into the weeds.)
- kick 7y agoEarlier this year the Media Lab did an absolutely ginormous automated transcription project. Off the top of my head, it was ~2.8 billion words. 13.1% error rate (vs. ~7% error rate for Google's proprietary solution). Found the paper on it: https://arxiv.org/pdf/1907.07073.pdf https://arxiv.org/pdf/1907.07073.pdf
- woodson 7y agoSadly they don’t (well, can’t) release the audio+transcripts as dataset, as they clearly don’t own the rights.
- kick 7y agoThey did release it as a dataset. I have a copy of it. It's massive. I'd recommend reading the paper, it has a link to a place where you can download it, and aside from that, it's also fascinating.
- woodson 7y agoIncluding the audio? I downloaded the transcripts from their bucket, but couldn’t find any information on how to obtain the corresponding recordings. At Interspeech 2019, the author basically told me that they couldn’t share it.
- woodson 7y agoIf you have access to the corresponding audio recordings, would you be able to share them?
- bgee 7y agoJust curious: $1/min sounds like quite a bit of money, are you paying for some professional to do this? If so have you compared with using Mechanical Turk?
- flurie 7y agoThat’s minute of recorded audio, and that’s a pretty standard transcription rate. Using anything less than a professional service will show in the quality of the output, and even many services don’t produce high quality transcripts, especially those that use temp (often undergraduate/graduate student) labor.
- bgee 7y agoMinutes of recorded audio make a lot of sense, thanks!