3 ms·
Thanks, I haven’t seen an easy and reliable way to do this using open source stuff yet. Theoretically just separating out speakers seems like it wouldn’t be tha
by eigenvalue 3y ago
Thanks, I haven’t seen an easy and reliable way to do this using open source stuff yet. Theoretically just separating out speakers seems like it wouldn’t be that hard; just compute a bunch of FFTs to arrive at a sort of frequency-based “voice fingerprint” for each speaker and then use something like XGboost to match up the audio for each second to one of the speakers. The problem is then what do you with that information? Turning those abstract speaker identifications into actual names would seem to require a fair bit of intelligence and picking up from contextual clues (like if the speaker identifies themselves or introduces another person by name). Anyway, I’ll look into it more. If it could be done reliably without overly complicating the setup, I agree that it would be useful.
- tehlike 3y agoThat part can be user input, if it needs to be. Sort of like post processing. Founder of Gladia shares some information on twitter, and i think there's some research papers you can find through that (my memory is fuzzy on when). IT's not a simple problem, especially when words include space fillers like "umm" etc. For me main use for such cases would be podcast. sometimes i just want to read them without listening.
- FanaHOVA 3y agoI built it for our podcast, I'm sure you can re-use it for YT as a source rather than raw audio file: https://github.com/FanaHOVA/smol-podcaster https://github.com/FanaHOVA/smol-podcaster
- jimmySixDOF 3y agoThe whisperX project has most of it covered they are integrating the new v3 model waiting for a ccp GGUF version also there is a workaround for the latest speaker diarization but it has an active user base working on it https://github.com/m-bain/whisperX https://github.com/m-bain/whisperX
- noman-land 3y agoI was dismayed to learn that this requires OpenAI API keys for speaker diarization.
- vsnf 3y agoDoes it? I only see references to HuggingFace api keys so it can do a one time download of some additional models.
- noman-land 3y agoOh, maybe that's what I saw. I'll have to look again. Thanks for keeping me honest.
- eigenvalue 3y agoThanks, took a look at it. Seems quite heavy though, lots of huge dependencies like pytorch and torchaudio, and seems like the speaker diarization requires a GPU if I'm not mistaken. And as another poster pointed out, it does require a Huggingface API key as well. I wanted to keep my script lighter weight and also GPU optional (i.e., a GPU will work and make it faster, but it also works acceptably with just the CPU). I really feel in my gut that the speaker diarization doesn't need to be so complicated or hard once you already have the accurate timestamps of each transcribed segment and the underlying audio file-- no reason why it shouldn't be able to run fine on a CPU and get good enough accuracy.
- huac 3y agoto zoom out a bit, this is known as the 'cocktail party problem' and is very much considered unsolved (particularly when there is an unknown number of speakers or overlapping speech).
- pavel_lishin 3y ago> Turning those abstract speaker identifications into actual names would seem to require a fair bit of intelligence As someone looking for this functionality, this is the easiest part for me to do manually. Just give me "Speaker A, Speaker B, Speaker C", and I can change their names. It's the breaking apart of audio into separate speakers that's difficult - I'm trying out a few different tools, and none of them do a great job so far.