3 ms·
> SOTA models are at ~5% WER for general speech. Do SOTA results measure performance in difficult environments though? Presumably dealing with people talking
by jeeeb 5y ago
> SOTA models are at ~5% WER for general speech.
Do SOTA results measure performance in difficult environments though?
Presumably dealing with people talking out of their car window next to a busy road would be a lot more difficult than dealing with a relatively clear audio recording.
EDIT: It looks like the linked results are for an audio book dataset. That seems like an optimal environment where you’re going to get clear enunciation with minimal background noise.