6 ms·
I’ve started doing this with ASR hypotheses from colloquial spontaneous speech. It tends to have similar issues. Lots of shady human ground truth especially w
by blackkettle 3y ago
I’ve started doing this with ASR hypotheses from colloquial spontaneous speech. It tends to have similar issues. Lots of shady human ground truth especially where addresses, alphanumeric sequences, repairs and repetitions and other essentially non read speech are concerned. The very large Whisper models are consistent in their transcription style and highly reliable as long as you pick strongly represented languages. And ChatGPT can do a very good job at comparing the linguistic coherence of hypotheses from multiple recognizers. Together these models can annotate, analyze and ingest far more data more consistently than human annotators at this point (at least in the best covered languages). We haven’t quite realized this as a community yet though, because the standard datasets we use for evaluation contain all these human inconsistencies. Wild times.