4 ms·
It's not as if people aren't trying to do that: https://github.com/openai/whisper/discussions/264 https://github.com/openai/whisper/discussions/264 I tried out
by coder543 4y ago
It's not as if people aren't trying to do that: https://github.com/openai/whisper/discussions/264 https://github.com/openai/whisper/discussions/264
I tried out this notebook about a month ago, and it was rough. After spending an evening improving it, I got everything "working", but pyannote was not reliable. I tried it against an hour-ish audio sample, and I found no way to tune pyannote to keep track of ~10 speakers over the course of that audio. It would identify some of the earlier speakers, but then it felt like it lost attention and would just start labeling every new speaker as the same speaker. There is an option to force the minimum number of speakers higher, and that just caused it to split some of the earlier speakers into multiple labels. It did nothing to address the latter half of the audio.
So, sure, someone should continue working on putting the pieces together, and I'm sure the notebook in the discussion I linked has probably improved since then, but I think pyannote itself needs some improvement first.
Sadly, I think using separate models for transcription and diarization ends up being clunky to the point that it won't ever be polished, no matter how good pyannote might get. If you have a podcast-like environment where people get excited and start talking over each other, then even if pyannote correctly identifies all of the speakers during the overlapping segments and when they spoke... Whisper cannot be used to separate speakers. You end up with either duplicate transcripts attributed to everyone involved, or something worse. Impressively, I have seen pyannote do exactly that, when it's working.
At the end of the day, I think someone is going to need to either train Whisper to also perform diarization, or we're going to need to wait until someone else open sources a model that does both transcription and diarization simultaneously. Unfortunately, it seems like most of these really big advances in ML only happen when a corporate benefactor is willing to dump money into the problem and then release the result, so we might be waiting awhile. I'm trying to learn more about machine learning, but I'm not at the point where I have any realistic chance of making such an improvement to Whisper. Maybe someone else around here can proven me wrong by just making it happen.
- password4321 4y agoSpeaker recognition is another piece that isn't usually as high a priority as recognizing the speech.
- mayeaux 4y agoIt's a new thing to me, I hadn't really considered it. Do they have that for movies and stuff? I can't think of a clear case when I've seen it