3 ms·
My experience is that what you're describing is only initial_prompt, and it only affects the first 30-second transcription window of the audio in question. It'
by coder543 3y ago
My experience is that what you're describing is only initial_prompt, and it only affects the first 30-second transcription window of the audio in question.
It's effectively useless for helping the model transcribe new words in longer content. That also wouldn't be a long-term solution anyways... no one wants to compile a huge list of "words Whisper probably doesn't know" and have to pass those in every time the model is being used. Even if that worked, it would also distort the transcription, since you're not saying you know which words are in the actual speech, you're just passing in a list of words. So, you could end up influencing Whisper to choose the wrong words, giving priority to this list of random words being passed in.
I am similarly curious about how we can train Whisper models to learn new words over time, unless OpenAI plans to release updated models themselves.
- regularfry 3y agoWould attention sinks work here? https://github.com/mit-han-lab/streaming-llm https://github.com/mit-han-lab/streaming-llm - it sounds like they might. In theory it doesn't involve retraining, it's just a change to how the data is managed between invocations.