3 ms·
It's a property of the model itself. > Input audio is split into 30-second chunks, converted into a log-Mel spectrogram, and then passed into an encoder. http
by bakkoting 3y ago
It's a property of the model itself.
> Input audio is split into 30-second chunks, converted into a log-Mel spectrogram, and then passed into an encoder.
https://openai.com/research/whisper https://openai.com/research/whisper