4 ms·
Transcription speed and accuracy keeps going up and it’s delightful to see the progress, I wish though more effort was dedicated to creating integrated solution
by pen2l 3y ago
Transcription speed and accuracy keeps going up and it’s delightful to see the progress, I wish though more effort was dedicated to creating integrated solutions that could accurately transcribe with speaker diarization.
- siraben 3y agoIs diarization only possible with stereo audio at the moment in whisper? If the voices aren't left/right split I don't think one can get them separated yet.
- pen2l 3y agoWhisper cannot perform speaker diarization at the moment, WhisperX and other solutions that can tend to use pyannot https://github.com/pyannote/pyannote-audio https://github.com/pyannote/pyannote-audio
- copypirate 3y agohttps://github.com/Wordcab/wordcab-transcribe https://github.com/Wordcab/wordcab-transcribe Author of Wordcab-Transcribe here. We use faster-whisper + NeMo for diarization, if you want to take a look.
- theossuary 3y agoNvidia's Nemo has some support for speaker recognition and speaker diarization too. https://docs.nvidia.com/deeplearning/nemo/user-guide/docs/en/v1.0.0/asr/speaker_diarization/intro.html https://docs.nvidia.com/deeplearning/nemo/user-guide/docs/en...
- atmosx 3y agoSame here. Whisper is really good at transcribing Greek but no diarization support, which makes it less than ideal for most use cases.
- userhacker 3y agoI'm the creator of Revoldiv.com, We do speaker diarization and transcription at the same time. Give it a try.