3 ms·
Using the large model, it works really well, even in low volume settings/speakers mumbling. Some of my transcripts are pharma related and Whisper stumbles on th
by elektor 3y ago
Using the large model, it works really well, even in low volume settings/speakers mumbling. Some of my transcripts are pharma related and Whisper stumbles on the drug names, but I’m pretty understanding of that.
- java_beyb 3y agowhat you're looking for is called diarization. almost all enterprise STTs do that, you can find individual libraries on GitHub too. fine-tuning whisper is a nightmare, I don't know what the interviews are for, but again most enterprise STTs offer customization. you can add medical terminology. ---Google, Amazon and Nuance have medical models but either expensive or not available for personal projects.
- elektor 3y agoThanks for that! Searching for diarization really helped me narrow down for what I was looking for.