4 ms·
Yeah, I will add post-processing to my prototype, too. I already prepared a detailed spec (prototyping new ways to do this, as well): https://github.com/Leftium
by Leftium 7mo ago
Yeah, I will add post-processing to my prototype, too. I already prepared a detailed spec (prototyping new ways to do this, as well): https://github.com/Leftium/rift-transcription/blob/main/specs/transforms.md https://github.com/Leftium/rift-transcription/blob/main/spec...
One idea I was tossing around was streaming transcription + batch re-transcription:
- Use streaming transcription, which works most of the time (for example, I've found the Web Speech API pretty good, as well as moonshine)
- If the streaming transcription was poor, select the bad part and re-transcribe with a more accurate batch transcription model.
- helro 7mo agoI tested something similar and continous re-transcription was the only way I could get close to batch-level accuracy. In my current implementation I’m fairly aggressive with it. I don’t rely much on streaming word confidence. Instead I continuously reprocess audio using a sliding window. As new audio comes in, it’s retranscribed together with the previous segment so the model always sees a longer context. That recovers a lot of the accuracy lost with streaming, but the amount of retranscription makes it hard to justify economically with cloud APIs. That’s why I’m focusing on a local-first approach for now.