3 ms·
Thanks, took a look at it. Seems quite heavy though, lots of huge dependencies like pytorch and torchaudio, and seems like the speaker diarization requires a GP
by eigenvalue 3y ago
Thanks, took a look at it. Seems quite heavy though, lots of huge dependencies like pytorch and torchaudio, and seems like the speaker diarization requires a GPU if I'm not mistaken. And as another poster pointed out, it does require a Huggingface API key as well.
I wanted to keep my script lighter weight and also GPU optional (i.e., a GPU will work and make it faster, but it also works acceptably with just the CPU). I really feel in my gut that the speaker diarization doesn't need to be so complicated or hard once you already have the accurate timestamps of each transcribed segment and the underlying audio file-- no reason why it shouldn't be able to run fine on a CPU and get good enough accuracy.