4 ms·
Ran it over our internal dataset of ~250 recordings of people saying british postcodes (all kinds of accents, etc) - it's competitive for sure! Soniox (stt-asy
by mnbbrown 6mo ago
Ran it over our internal dataset of ~250 recordings of people saying british postcodes (all kinds of accents, etc) - it's competitive for sure!
Soniox (stt-async-v4): 176/248 (71.0%)
ElevenLabs (scribe_v2): 170/248 (68.5%)
AssemblyAI (universal-3-pro): 166/248 (66.9%)
Deepgram (nova-3): 158/248 (63.7%)
AssemblyAI (universal-2): 148/248 (59.7%)
Cohere (transcribe-03-2026): 148/248 (59.7%)
Speechmatics (enhanced): 134/248 (54.0%)
P.s. how do I get this to render correctly on here?
- Bolwin 6mo agoTry two newlines between each one
- ChrisMarshallNY 6mo agoThat, or add 4 spaces before each line (renders as a <pre>).
- mkl 6mo agoTwo spaces: https://news.ycombinator.com/formatdoc https://news.ycombinator.com/formatdoc It's for code though, not lists or bullet points.
- jilijeanlouis 6mo agodid you try gladia: ranking #1 on STT blind test https://compare-stt.com/ https://compare-stt.com/
- scotty79 6mo agoThis benchmark should have Whisper large-v3 as one of the models.
- mnbbrown 6mo agoAdded gladia.. - 1. Soniox (stt-async-v4): +176 new cases, running total 176/248 (71.0%) - 2. ElevenLabs (scribe_v2): +26 new cases, running total 202/248 (81.5%) - 3. Speechmatics (enhanced): +12 new cases, running total 214/248 (86.3%) - 4. NVIDIA Parakeet (TDT 0.6B v2): +6 new cases, running total 220/248 (88.7%) - 5. Mistral (voxtral-mini): +3 new cases, running total 223/248 (89.9%) - 6. Gladia: +2 new cases, running total 225/248 (90.7%) - 7. AssemblyAI (universal-2): +1 new cases, running total 226/248 (91.1%) - 8. Deepgram (nova-3): +1 new cases, running total 227/248 (91.5%) - 9. Cohere (transcribe-03-2026): +0 new cases, running total 227/248 (91.5%) - 10. AssemblyAI (universal-3-pro): +0 new cases, running total 227/248 (91.5%)
- yorwba 6mo agoIs the human baseline 248/248?
- walthamstow 6mo agoAssuming all the accents are British, I doubt it. I probably couldn't get all 248 myself.
- mnbbrown 6mo agoThey are all transcribed by multiple blinded "accent natives". But yes, your point is valid - going to see if I can tease out the "single person accuracy".