3 ms·
Whisper has the encoder-decoder architecture, so it's hard to run streaming efficiently, though whisper-streaming is a thing. https://kyutai.org/next/stt https
by nomad_horse 1y ago
Whisper has the encoder-decoder architecture, so it's hard to run streaming efficiently, though whisper-streaming is a thing.
https://kyutai.org/next/stt https://kyutai.org/next/stt is natively streaming STT.
- woodson 1y agoThere are many streaming ASR models based on CTC or RNNT. Look for example at sherpa (https://github.com/k2-fsa/sherpa-onnx https://github.com/k2-fsa/sherpa-onnx), which can run streaming ASR, VAD, diarization, and many more.