4 ms·
I have a pet ML project that I am doing for fun. I am trying to build a custom transcription and diarizer model for a friend's podcast[1]. My initial solution i
by SubiculumCode 2y ago
I have a pet ML project that I am doing for fun. I am trying to build a custom transcription and diarizer model for a friend's podcast[1]. My initial solution involved a straight forward implementation using Whisper medium for transcription, and Nemo for diarizing, based on [1]. The results are not bad generally, but since my application involves a fixed set of five known speakers, I thought surely I could fine tune the nemo (or pyannote) diarizer model on their voices to improve accuracy.
Audio samples are easily obtained from their podcast, but manual data labeling is painful for a hobby activity. Further, from what I understand, the real difficulty in performant diarizer models is not speaker recognition generally, but specifically speaker recognition while there is overlapping speech between multiple speakers. I am not even sure how to best implement a labeling procedure for segments with overlapping speech.
I started to wonder whether I might bootstrap a decent sample by leveraging TTS vocal cloning models to simulate the five speakers in dialogues with overlapping speech segments. So I ask HN, is this hopelessly naive, or potentially useful technique? Also, any other advice?
[1] https://www.3d6downtheline.com/ https://www.3d6downtheline.com/ [2] https://github.com/MahmoudAshraf97/whisper-diarization/ https://github.com/MahmoudAshraf97/whisper-diarization/
- tarasglek 2y agoUnclear from docs, does your solution support inferring number of speakers from audio? Found it a bit frustrating that this wasn't automatic in diarization algos I tried last year
- SubiculumCode 2y agoThe solution that this GitHub provided automatically determines the number of speaker labels, but will often create extra speaker classes for a few exerpts in the stream. You can prespecify the number of speakers I believe for better performance.