4 ms·
I’ve tested at least 20 STT models in a benchmark I’ve set up with German, Italian and English voices from meetings in my company. The voices contained very in
by Lucasoato 1mo ago
I’ve tested at least 20 STT models in a benchmark I’ve set up with German, Italian and English voices from meetings in my company.
The voices contained very industry specific words, the languages changed from one sentence to another, sometimes words in a language were mentioned while a discussion was in another.
The only local model that satisfies me is Voxtral Mini 3b, the only paid API that is slightly better is eleven labs. Yes, Voxtral might not reach the best score in the benchmarks, but to me, it just solves a problem. It might not be the best in terms of speed... but that’s not a problem for me.
Happy to test this new model from Google but I’m not sure I’d go with that instead of something that can run so easily in my machine.
- cnxhk 1mo agoAny possibility to share some of the eval audio?
- Lucasoato 1mo agoOf course not, but it’s easily replicable just by mixing different languages conversations together, adding a word here and there of some very specific German jargon.
- hinnisdael 1mo agoAgree about Voxtral being the only model — local or cloud — that handles multilingual conversation really well. I‘m not sure what they do differently, but mixed-language sentences and industry terms don‘t seem to faze it where other model begin to struggle.
- kridsdale1 1mo agoI use Parakeet 3. How does that fare in your testing?
- Lucasoato 1mo agoI love it but it misses the business specific words when in different language. Sometimes it stretches them out to fit unrelated words in the language of the rest of the conversation. I miss its speed though.
- wahnfrieden 1mo agoTry the recent MOSS one? It’s very good
- crossroadsguy 1mo agoDo you use it on a desktop? Mac by any chance? What's your setup? I've been looking to find a simple and fast dictation app for English but almost everything I've tried (from Handy to many apps, eg some with Whisper in their names, after the model I assume) just don't work well. Apple's offering is worse than those though. I even tried with local enhancement models.
- pomtato 1mo agoYou should give VoiceInk[1] a go. It works great and has decent latency. https://github.com/Beingpax/VoiceInk https://github.com/Beingpax/VoiceInk
- brmkr 1mo agoI am biased as a developer on the project but you should give Epilude [1] a go if you want local dictation on a Mac. We’ve fine-tuned open-weight models to make them better (in our benchmarks) at cleaning up and formatting what you say so you don’t have to edit what you dictate. https://epilude.com https://epilude.com
- jwr 1mo agoI used MacWhisper for a long time and then recently switched to the free TypeWhisper. Both work very well on a Mac.
- stavros 1mo agoOut of curiosity, what was wrong with Handy? I use it and it works fine.
- LeBit 1mo agoI use Handy all the time and it is just perfect. I guess OP didn’t chose the right STT model.
- deleted 1mo ago[deleted]
- 1mo ago
- anton000 1mo agohave you tried groq whisper-turbo? works well for me for multiple different languages
- orbital-decay 1mo agoThis problem is legitimately hard and needs high cognitive abilities. Even the biggest generalist models struggle with memes and lingo salad that sound immediately intuitive for an out-of-the-loop human, and I'm talking about text comprehension. Modern models are optimized for decision making and are worse in that than old ones optimized for creative writing, but those also struggled. You should probably not expect a small STT model like Voxtral or Parakeet to do any better, unless it's laser focused on that area in particular and sucks in everything else.
- totetsu 1mo agoWhat i want is a model that outputs its predictions and their scores along with the text it choose. So I could flag something it's getting consistently wrong, like mis-predicting a technical term or name, or acronym, and say replace it with my correction, and have the ui be able to smartly replace that in the whole text so far and future parts. even better would be the ability to feed this back into the model for future runs.
- Lucasoato 1mo agoGood news: most of these models can include a prompt that steers the transcription; if you use frequently a word you just invented, add it there and it will be transcribed correctly more likely.
- jwr 1mo agoIt's interesting that the results can be so different depending on the person, the use case and even the microphone used. I use dictation a lot, so I try to stay up to date with the latest models as much as I can. So far, for my needs, nothing could beat Whisper Large v3. I keep hearing that Nvidia Parakeet models are better, but they just don't work as well for me, even though they are unquestionably faster. Things change a lot if you need to speak to the model in multiple languages. There are very few models out there that can automatically detect the language spoken and produce correct output. I tried the larger Voxtral models, but they didn't work for me at all. When I spoke to them in Polish, they produced output in Russian or Ukrainian. For now, I settled on creating my own plugin for TypeWhisper, which runs Whisper Large on the GPU and does it much faster than pretty much anything else out there. But I'm still hoping that something better will come along, as Whisper Large is quite old at this point.
- Melatonic 1mo agoI wonder if cross training models specifically on people who combine languages (like Spanglish) would help with this. Surely must be patterns in what words people choose to use in each language
- Nitrolo 1mo agoTo be fair, Mistral doesn't list Ukrainian as one of the supported languages.
- jwr 1mo agoI spoke English and Polish to it. It produced Russian or Ukrainian as output.
- HyperAI 1mo ago[dead]
- PeriPan 1mo agoI suggest try openwhisperer.com, open-source. DMG or build from source (GitHub) Transcribing is perfect in many language and, it focuses on a selected app. Also it does TTS. Using a combi of the large Whisper and koroko/supersonic for TTS. Plugs into dev environments via MCP and hooks (Claude, Codex).
- Computer0 1mo agoI am also using voxtral for local overnight batch runs and voxtral via api for more expedient use cases.
- fumeux_fume 1mo agoI also found eleven labs gave me the best results when transcribing French dialog from a 60s television show. Voxtral's output format was the easiest to work with when attempting to create actual subtitles.
- jiehong 1mo agoVoxtral nomenclature is a bit confusing. I never remember which is newer for exemple (the number of parameters isn’t always helping).