4 ms·
Any plans to directly support diarization or voiceprinting?
by staticautomatic 2y ago
Any plans to directly support diarization or voiceprinting?
- jeffharris 2y agoWe're thinking about diarization (adding time awareness to GPT models) but no firm plans to share just yet
- simonw 2y agoThe feature I want is speaker differentiation - I want to feed in an audio file and get back a transcript with "Speaker 1: ..., Speaker 2: ..." indications. That plus timestamps would be incredible. The Google Gemini 2.0 models are showing some promise with this, I can't speak to their reliability just yet though.
- runeb 2y agoI had good results with pyannote and the following model for that use case in the past https://huggingface.co/pyannote/speaker-diarization-3.1 https://huggingface.co/pyannote/speaker-diarization-3.1
- infecto 2y agoI thought Deepgram already did speaker diarization (which is differentiation) pretty well. That and it can include timestamps plus other metadata.
- thot_experiment 2y agoWhisperX does all of this, I use it all the time to transcribe meeting notes. Both speaker differentiation and individual word timestamps.
- youssefabdelm 2y agoJeff you know what would be magical? Not just vanilla diarization "Speaker 1" and "2" but if the model can know from the conversation this speaker was referred to as "Jeff Harris" or "Jeff" so it uses that instead.
- youssefabdelm 2y agoOr if we could even provide samples of what an example speaker sounds like in general so that it would always classify them the way we want.