4 ms·
I mean, Microsoft claimed in 2016 to surpass human performance and google subsequently claimed in 2017, and yet anyone that regularly uses voice dictation knows
by tbenst 6y ago
I mean, Microsoft claimed in 2016 to surpass human performance and google subsequently claimed in 2017, and yet anyone that regularly uses voice dictation knows that <5% WER is far from realistic due to lack of generalization to new audio/noise regimes and domain-specific vocabulary.
- yorwba 6y ago5% WER or less isn't that unrealistic. It's just not very good. It means that there's an error every 20 words on average, which with a slow rate of speech of 2 words per second amounts to an error every ten seconds. That humans can still communicate effectively with each other despite a higher transcription error rate is just due to the the way humans do transcriptions: listen to a section of audio, then write down what they heard. If they understood the meaning, they'll likely write down a transcription that retains this meaning, but they might drop an article or duplicate it or switch the order of two words. It's easier for an algorithm to avoid those mistakes of inattention, but harder to prevent errors that change the meaning. So WER isn't a perfect indicator of transcription quality. It also matters which words are affected. I wouldn't be surprised if people preferred a human-written transcription that reads well despite a high WER over a machine-generated one with fewer errors, but more obvious ones. Bonus: find the word duplication in this comment.
- candiodari 6y agoThat would be easy to duplicate, wouldn't it? Train GPT-2 (because no API key required) to just repeat the sentence correctly, feed the output of the transcription through. I'm not sure if these models do that.
- darepublic 6y agoYeah have a model predict the misheard word in a sentence where you replace one word with a speech recognition error
- Aerroon 6y ago>That humans can still communicate effectively with each other despite a higher transcription error rate is just due to the the way humans do transcriptions: listen to a section of audio, then write down what they heard. Part of it is that human communication has something similar to error correction built in. We recognize patterns of words and can guess what the next word is based on that. This allows us to discern the meaning without always understanding every word. We effectively fill in the blanks.
- bumbledraven 6y ago> Bonus: find the word duplication in this comment. Gur ercrngrq jbeq vf gur. V unq gb hfr n cebtenz gb svaq vg: crey -yar 'juvyr (/\o(\j+)\o\f+\o\1\o/t) {cevag $1}'