3 ms·
I tested something similar and continous re-transcription was the only way I could get close to batch-level accuracy. In my current implementation I’m fairly a
by helro 7mo ago
I tested something similar and continous re-transcription was the only way I could get close to batch-level accuracy.
In my current implementation I’m fairly aggressive with it. I don’t rely much on streaming word confidence. Instead I continuously reprocess audio using a sliding window. As new audio comes in, it’s retranscribed together with the previous segment so the model always sees a longer context.
That recovers a lot of the accuracy lost with streaming, but the amount of retranscription makes it hard to justify economically with cloud APIs. That’s why I’m focusing on a local-first approach for now.