3 ms·
Answering your question from the small end: I picked my STT by testing self-correction handling. My tool cleans up spoken drafts, and the failure that mattered
by jobuildsstuff 2mo ago
Answering your question from the small end: I picked my STT by testing self-correction handling. My tool cleans up spoken drafts, and the failure that mattered wasn't word accuracy, it was "meet Tuesday, no wait, Wednesday": a raw transcript of that is worse than useless, and models differ a lot in how gracefully downstream cleanup can recover. Your spontaneous-speech testing sounds close to this already. Do the boards score corrections and disfluencies specifically, or do they fold into overall accuracy?
- taiuo_ops 2mo agoWe hit the same problem building a voice-memo-to-notes tool: "meet Tuesday, no wait, Wednesday" is exactly the failure mode. We ended up not trying to fix it at the ASR layer at all. We run faster-whisper (large-v3, local) and let the raw transcript keep the self-correction intact, including the discarded "Tuesday." The summarization pass resolves it, because it has document-level context the transcriber doesn't, and it's cheap to re-run on a short transcript. A technically "wrong" transcript still produces the right structured output most of the time. The failure mode shifts rather than disappears though: on longer memos where corrections stack up, the summarizer occasionally keeps the wrong one, and that's much harder to catch than a raw WER miss because the sentence still reads fine. Do you know of anyone scoring disfluency-recovery as its own metric instead of folding it into WER? Feels like exactly the kind of thing that would need a downstream-task-aware eval rather than a transcript-only one.