3 ms·
I recently tested a whole bunch of models (basically whatever was recent and available on openrouter filtering by audio input) for STT for mixed language audio
by sireat 1mo ago
I recently tested a whole bunch of models (basically whatever was recent and available on openrouter filtering by audio input) for STT for mixed language audio - mostly English mixed with Latvian.
In the end I was basically forced to go with Scribe (v2) it was only one that had consistently high quality across multilingual speech with multiple speakers .
Crucially it correctly identified multiple speakers across hour of audio.
I would love to go with something like Transcribe if it gets close to this type of performance.