4 ms·
Even in the commercial space, there’s a lack of production grade ASR APIs that support diarization and word level timestamps. My experiences with Google’s Chir
by bartman 6mo ago
Even in the commercial space, there’s a lack of production grade ASR APIs that support diarization and word level timestamps.
My experiences with Google’s Chirp have been horrendous, with it sometimes skipping sections of speech entirely, hallucinating speech where the audio contains noise, and unreliable word level timestamps. And this all is even with using their new audio prefiltering feature.
AWS works slightly better, but also has trouble with keeping word level timestamps in sync.
Whisper is nice but hallucinates regularly.
OpenAI’s new transcription models are delivering accurate output but do not support word level timestamps…
A lot of this could be worked around by sending the resulting transcripts through a few layers of post processing, but… I just want to pay for an API that is reliable and saves me from doing all that work.
- stavros 6mo agoIsn't Elevenlabs the best in this?
- bartman 6mo agoI've not tested their speech-to-text yet, but based on the docs it looks promising. Thanks for the suggestion!
- stavros 6mo agoIt's fantastic, and their diarization is spot on as well.
- gardnr 6mo agoThey can have issues with the timestamps: https://github.com/elevenlabs/elevenlabs-python/issues/707 https://github.com/elevenlabs/elevenlabs-python/issues/707
- catlifeonmars 6mo agoI wonder if you could run multiple models and average out the timestamps, kind of like how atomic clocks are used together and not separately