5 ms·
Show HN: Moonshine Open-Weights STT models – higher accuracy than WhisperLargev3
I wanted to share our new speech to text model, and the library to use them effectively. We're a small startup (six people, sub-$100k monthly GPU budget) so I'm proud of the work the team has done to create streaming STT models with lower word-error rates than OpenAI's largest Whisper model. Admittedly Large v3 is a couple of years old, but we're near the top the HF OpenASR leaderboard, even up against Nvidia's Parakeet family. Anyway, I'd love to get feedback on the models and software, and hear about what people might build with it.
- regularfry 7mo agoOh this is fantastic. I'm most interested to see if this reaches down to the raspberry pi zero 2, because that's a whole new ballgame if it does.
- cyanydeez 7mo agoNo LICENSE no go
- bangaladore 7mo agoThere is a license blurb in the readme. > This code, apart from the source in core/third-party, is licensed under the MIT License, see LICENSE in this repository. > The English-language models are also released under the MIT License. Models for other languages are released under the Moonshine Community License, which is a non-commercial license. > The code in core/third-party is licensed according to the terms of the open source projects it originates from, with details in a LICENSE file in each subfolder.
- mkl 7mo agoThe LICENSE file that refers to is missing. There's one in the python folder, but not for the rest of the code.
- namibj 7mo agoIANAL. Presuming (I haven't checked myself) the git author information supports this, it should be fine to treat this as licensing the code it specifies under MIT; based on that license name being (to my understanding) unambiguous and license application being based on contract law and contract law basically having at it's very core the principle of "meeting of the minds" along with wilful infringement being really really hard to even argue for if the only thing that's separating it from being 100% clearly licensed in all proper ways being not copying in an MIT `LICENSE` template with date and author name pasted into it.
- deleted 7mo ago[deleted]
- altruios 7mo agoreading through readme.md "License This code, apart from the source in core/third-party, is licensed under the MIT License, see LICENSE in this repository. The English-language models are also released under the MIT License. Models for other languages are released under the Moonshine Community License, which is a non-commercial license. The code in core/third-party is licensed according to the terms of the open source projects it originates from, with details in a LICENSE file in each subfolder."
- lostmsu 7mo agoHow does it compare to Microsoft VibeVoice ASR https://news.ycombinator.com/item?id=46732776 https://news.ycombinator.com/item?id=46732776 ?
- armcat 7mo agoThis is awesome, well done guys, I’m gonna try it as my ASR component on the local voice assistant I’ve been building https://github.com/acatovic/ova https://github.com/acatovic/ova. The tiny streaming latencies you show look insane
- ac29 7mo agoNo idea why 'sudo pip install --break-system-packages moonshine-voice' is the recommended way to install on raspi? The authors do acknowledge this though and give a slightly too complex way to do this with uv in an example project (FYI, you dont need to source anything if you use uv run)
- deleted 7mo ago[deleted]
- g-mork 7mo agoHow does this compare to Parakeet, which runs wonderfully on CPU?
- pzo 7mo agohaven't tested yet but I'm wondering how it will behave when talking about many IT jargon and tech acronyms. For those reason I had to mostly run LLM after STT but that was slowing done parakeet inference. Otherwise had problems to detect properly sometimes when talking about e.g. about CoreML, int8, fp16, half float, ARKit, AVFoundation, ONNX etc.
- sroussey 7mo agoonnx models for browser possible?
- asqueella 7mo agoFor those wondering about the language support, currently English, Arabic, Japanese, Korean, Mandarin, Spanish, Ukrainian, Vietnamese are available (most in Base size = 58M params)
- Karrot_Kream 7mo agoAccording to the OpenASR Leaderboard [1], looks like Parakeet V2/V3 and Canary-Qwen (a Qwen finetune) handily beat Moonshine. All 3 models are open, but Parakeet is the smallest of the 3. I use Parakeet V3 with Handy and it works great locally for me. [1]: https://huggingface.co/spaces/hf-audio/open_asr_leaderboard https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
- reitzensteinm 7mo agoParakeet V3 is over twice the parameter count of Moonshine Medium (600m vs 245m), so it's not an apples to apples comparison. I'm actually a little surprised they haven't added model size to that chart.
- agentifysh 7mo agoSo I'm kinda new to this whole parakeet and moonshine stuff, and I'm able to run parakeet on a low end CPU without issues, so I'm curious as to how much that extra savings on parameters is actually gonna translate. Oh and I type this in handy with just my voice and parakeet version three, which is absolutely crazy.
- bytesandbits 7mo agoparakeet v3 has a much better RTFx than moonshine, it's not just about parameter numbers. Runs faster. https://huggingface.co/spaces/hf-audio/open_asr_leaderboard https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
- regularfry 7mo agoIt is about the parameter numbers if what you care about is edge devices with limited RAM. Beyond a certain size your model just doesn't fit, it doesn't matter how good it is - you still can't run it.
- bytesandbits 7mo agoI am not sure what "edge" device you want to run this on, but you can compress parakeet to under 500MB on RAM / disk with dynamic quants on-the-fly dequantization (GGUF or CoreML centroid palettization style). And retain essentially almost all accuracy. And just to be clear, 500MB is even enough for a raspberry Pi. Then your problem is not memory, is FLOPS. It might run real-time in a RPi 5, since it has around 50 GFLOPS of FP32, i.e. 100 GFLOPS of FP16. So about 20-50 times less than a modern iPhone. I don't think it will be able to keep it real time, TBF, but close. regardless, this model with such quantization strategy runs real time at +10x real-time factor even in 6-year old iPhones (which you can acquire for under $200) and offline at a reasonable speed, essentially anywhere. You get the best of both worlds: the accuracy of a whisper transformer at the speed and footprint of a small model.
- aplomb1026 7mo ago[flagged]
- rob 7mo agoAnother fake bot account. This is getting ridiculous. It's every other thread here now. Timestamp 1: 2026-02-25T00:31:28 1771979488 https://news.ycombinator.com/item?id=47145661 https://news.ycombinator.com/item?id=47145661 Timestamp 2: 2026-02-25T00:32:03 1771979523 https://news.ycombinator.com/item?id=47145666 https://news.ycombinator.com/item?id=47145666 Two detailed large comments in two different threads in a 35 second span from a new account.
- nmstoker 7mo agoAny plans regarding JavaScript support in the browser? There was an issue with a demo but it's missing now. I can't recall for sure but I think I got it working locally myself too but then found it broke unexpectedly and I didn't manage to find out why.
- Leftium 7mo agoWASM-based port: https://github.com/moonshine-ai/moonshine-js https://github.com/moonshine-ai/moonshine-js I also did a survey of other in-browser transcription solutions: https://github.com/Leftium/rift-transcription/blob/main/reference/streaming-transcription-demos.md#in-browser--wasm-offline-no-server https://github.com/Leftium/rift-transcription/blob/main/refe... - Notably, there is an (unrelated?) moonshine demo based on transformers.js (using WebGPU) with WASM fallback.
- alexnewman 7mo agoIf only it did Doric
- 999900000999 7mo agoVery cool. Anyway to run this in Web assembly, I have a project in mind
- fareesh 7mo agoAccuracy is often presumed to be english, which is fine, but it's a vague thing to say "higher" because does it mean higher in English only? Higher in some subset of languages? Which ones? The minimum useful data for this stuff is a small table of language | WER for dataset
- saltwounds 7mo agoStreaming transcription is crazy fast on an M1. Would be great to use this as a local option versus Wispr Flow.
- francislavoie 7mo agoI've helped many Twitch streamers set up https://github.com/royshil/obs-localvocal https://github.com/royshil/obs-localvocal to plug transcription & translation into their streams, mainly for German audio to English subtitles. I'd love a faster and more accurate option than Whisper, but streamers need something off-the-shelf they can install in their pipeline, like an OBS plugin which can just grab the audio from their OBS audio sources. I see a couple obvious problems: this doesn't seem to support translation which is unfortunate, that's pretty key for this usecase. Also it only supports one language at a time, which is problematic with how streamers will frequently code-switch while talking to their chat in different languages or on Discord with their gameplay partners. Maybe such a plugin would be able to detect which language is spoken and route to one or the other model as needed?
- mattmcegg 7mo agoI released a OBS plugin (and optional RTMP relay) that does exactly this. It can do real time translated captions and voice cloning/dubbing. The plugin lets you choose an audio source, then creates each language's captions and dub as new Sources. Use them however you'd like! check it out! https://streamfluent.ai https://streamfluent.ai
- starkparker 7mo agoImplemented this to transcribe voice chat in a project and the streaming accuracy in English on this was unusable, even with the medium streaming model.
- heftykoo 7mo agoClaiming higher accuracy than Whisper Large v3 is a bold opening move. Does your evaluation account for Whisper's notorious hallucination loops during silences (the classic 'Thank you for watching!'), or is this purely based on WER on clean datasets? Also, what's the VRAM footprint for edge deployments? If it fits on a standard 8GB Mac without quantization tricks, this is huge.
- guerython 7mo ago[flagged]
- PranayKumarJain 7mo ago[flagged]
- regularfry 7mo agoTangentially, have you got any idea what the equivalent "partial tokens revised" rate for humans is? I know I've consciously experienced backtracking and re-interpreting words before, and presumably it happens subconsciously all the time. But that means there's a bound on how low it's reasonable to expect that rate to be, and I don't have an intuition for what it is.
- oezi 7mo agoDo you also support timestamps the detected word or even down to characters?
- raybb 7mo agofyi the typepad link in your bio is broken
- deleted 7mo ago[deleted]
- RobotToaster 7mo ago> Models for other languages are released under the Moonshine Community License, which is a non-commercial license. Weird to only release English as open weights.
- riedel 7mo agoI find it an even more weird practice for anyone working with speech or text models not in the first paragraph name the language it is meant for (and I do not mean the programming language bindings). How many English native speakers are there 5% of the world population?
- RobotToaster 7mo agoApproximately yes, although another 15% are non-native English speakers. Chinese is a close second for total speakers.
- dagss 7mo agoVery exciting stuff! hear about what people might build with it My startup is making software for firefighters to use during missions on tablets, excited to see (when I get the time) if we can use this as a keyboard alternative on the device. It's a use case where avoiding "clunky" is important and a perfect usecase for speech-to-text. Due to the sector being increasingly worried about "hybrid threats" we try to rely on the cloud as little as possible and run things either on device or with the possibility of being self-hosted/on-premise. I really like the direction your company is going in in this respect. We'd probably need custom training -- we need Norwegian, and there's some lingo, e.g., "bravo one two" should become "B-1.2". While that can perhaps also be done with simple post-processing rules, we would also probably want such examples in training for improved recognition? Have no VC funding, but looking forward to getting some income so that we can send some of it in your direction :)
- steinvakt2 7mo agoInteresting. Can we get in touch? I just sold my webapp/saas where I used NB-Whisper to transcribe Norwegian media (podcast, radio, TV) and offer alerts and search by indexing it using elasticsearch. Edit: It was https://muninai.eu https://muninai.eu (I shut down the backend server yesterday so the functionality is disabled).
- dagss 7mo agoSure! I didn't find your contact info but drop me an email at dag@syncmap.no.
- Ross00781 7mo agoThe streaming architecture looks really promising for edge deployments. One thing I'm curious about: how does the caching mechanism handle multiple concurrent audio streams? For example, in a meeting transcription scenario with 4-5 speakers, would each stream maintain its own cache, or is there shared state that could create bottlenecks?
- binome 7mo agoI vibe-trained moonshine-tiny on amateur radio morse code last weekend, and was surprised at the ~2% CER I was seeing in evals and over the air performance was pretty acceptable for a couple hour run on a 4090.
- sourcetms 7mo agoI'm offering support for this in Resonant - Already set up and running this week. It's incredible for a live transcription stream - the latency is WOW. https://www.onresonant.com/ https://www.onresonant.com/ For the open source folks, that's also set up in handy, I think.
- admiralrohan 7mo agoIs this alternative to Whispr Flow?
- nivcmo 7mo ago[dead]
- devcraft_ai 7mo ago[flagged]
- T0mSIlver 7mo agoCongrats on the results. The streaming aspect is what I find most exciting here. I built a macOS dictation app (https://github.com/T0mSIlver/localvoxtral https://github.com/T0mSIlver/localvoxtral) on top of Voxtral Realtime, and the UX difference between streaming and offline STT is night and day. Words appearing while you're still talking completely changes the feedback loop. You catch errors in real time, you can adjust what you're saying mid-sentence, and the whole thing feels more natural. Going back to "record then wait" feels broken after that. Curious how Moonshine's streaming latency compares in practice. Do you have numbers on time-to-first-token for the streaming mode? And on the serving side, do any of the integration options expose an OpenAI Realtime-compatible WebSocket endpoint?
- Leftium 7mo agoMy app uses this moonshine-voice python package, so you can experience it yourself here: https://rift-transcription.vercel.app/local-setup https://rift-transcription.vercel.app/local-setup I made moonshine the default because it has the best accuracy/latency (aside from Web Speech API, but that is not fully local) I plan to add objective benchmarks in the future, so multiple models can be compared against the same audio data... --- I made a custom WebSocket server for my project. It defines its own API (modeled on the Sherpa-onnx API), but you could adjust it to output the OpenAI Realtime API: https://github.com/Leftium/rift-local https://github.com/Leftium/rift-local (note rift-local is optimized for single connections, or rather not optimized to handle multiple WS connections)
- dSebastien 7mo agoI've been using Moonshine since V1 and the results are really great. I'd say on par with Parakeet V3 while working really well with CPU only.
- fittingopposite 7mo agoWhich program does support it to allow streaming? Currently using spokenly and parakeet but would like to transition to a model that is streaming instead of transcribing chunk wise.
- fudged71 7mo agoIf it's using ONNX, can this be ported to Transformers.js?
- Ross00781 7mo agoOpen-weight STT models hitting production-grade accuracy is huge for privacy-sensitive deployments. Whisper was already impressive, but having competitive alternatives means we're not locked into a single model family. The real test will be multilingual performance and edge device efficiency—has anyone benchmarked this on M-series or Jetson?
- Leftium 7mo agoTry Moonshine with a browser GUI: uv tool install rift-local && rift-local serve --open This opens RIFT[1], my web frontend for local transcription with a copy button. You can also compare against Web Speech API and other models (including cloud API's). https://github.com/Leftium/rift-local https://github.com/Leftium/rift-local [1]: https://rift-transcription.vercel.app/local-setup https://rift-transcription.vercel.app/local-setup
- Paddyz 7mo ago[dead]