5 ms·
Parakeet.cpp – Parakeet ASR inference in pure C++ with Metal GPU acceleration
- noahkay13 7mo agoI built a C++ inference engine for NVIDIA's Parakeet speech recognition models using Axiom(https://github.com/Frikallo/axiom https://github.com/Frikallo/axiom) my tensor library. What it does: - Runs 7 model families: offline transcription (CTC, RNNT, TDT, TDT-CTC), streaming (EOU, Nemotron), and speaker diarization (Sortformer) - Word-level timestamps - Streaming transcription from microphone input - Speaker diarization detecting up to 4 speakers
- aaronbrethorst 7mo agoI see a number of references to macOS support in your docs for Axiom. Can this run on iOS?
- noahkay13 7mo agoTheoretically, yes? This hasent been tested but xcode has great c++ interop and the goal with Axiom and now parakeet.cpp is to be used for portable deployments so making that process easier is definitely on the roadmap.
- computerex 7mo agoOh hey I just implemented this in golang. Mine implementation heavily optimized for cpu.
- pdyc 7mo agocan you share your repo.
- ghostpepper 7mo agoOff topic but if anyone is looking for a nice web-GUI frontend for a locally-hosted transcription engine, Scriberr is nice https://github.com/rishikanthc/Scriberr https://github.com/rishikanthc/Scriberr
- deleted 7mo ago[deleted]
- nullandvoid 7mo agoI've been using handy with parakeet on both Windows and mac, and have been very impressed. Hoe does this compare?
- jack_pp 7mo agoHandy only supports microphone input not files afaik
- qwertox 7mo agoI was also impressed with Handy. I played around with it this week, and when you enable advanced mode and add a post-transcription AI model to point to your own server which mimics a minimal ChatGPT-compatible behavior, then you can use it to modify the output, even return an empty string if you noticed that the transcript was more targeted to do other stuff ("turn the lights on"), if you then return an empty string, it won't inject keypresses. So one gets the best for both worlds: transcription for dictation and transcription to trigger events. If I now only could let it listen constantly and react to voice, so that no push to talk is active, that would be nice. Maybe this project here could be used for that. Also, this seems to support streaming transcription.
- potatoman22 7mo agoHow did you install parakeet? It was a nightmare to install on windows
- nullandvoid 7mo agoSorry for the late reply - it just worked out the box via the GUI for me using Handy. Win 11 with 4070ti, 7600x
- antirez 7mo agoRelated: https://github.com/antirez/qwen-asr https://github.com/antirez/qwen-asr https://github.com/antirez/voxtral.c https://github.com/antirez/voxtral.c Qwen-asr can easily transcribe live radio (see README) in any random laptop. It looks like we are going to see really cool things on local inference, now that automatic programming makes a lot simpler to create solid pipelines for new models in C, C++, Rust, ..., in a matter of hours.
- pjmlp 7mo agoWhich is why long term current programming languages will eventually become less relevant in the whole programming stack, as in get the computer to automate tasks, regardless how.
- FpUser 7mo agoAssuming RAM prices will not make it totally unaffordable. Current situation is atrocious and big infrastructure corps seem to love it, they do not want independent computing. Alternatively they might build specialized branded hardware which people could only use for what corps allow them to do for nice monthly fee. Another problem is too much abstraction on input spec level. The other day I asked Claude to generate few classes. When reviewing the code I noticed it doing full scan for ranges on one giant set. This would bring my backend to a halt. After pointing it out to Claude it had smartened up to start with lower_bound() call. When there are no people to notice such things what do you think we are going to have?
- pjmlp 7mo agoAgreed, in regards to prices, it appears to be the new gold, lets see how this gets sorted out, with NPUs, FPGAs, analog (Cerebas),... Now the abstraction I am with you on that, I foresee a more formal way to give specifications, but more suitable for natural language as input, or even proper mathematics, than the languages we have been using thus far. Naturally we aren't there yet.
- 7mo ago
- rowanG077 7mo agoIs there anything truly low latency(sub 100ms)? Speech recognition is so cool but I want it to be low latency.
- moffkalast 7mo agoParakeet does streaming I think, so if you throw enough compute at it, it should be. The closest competitor is whisper v3 which is relatively slow, maybe Voxtral but it's still very new.
- regularfry 7mo agoThere's a minimum possible latency just given the structure of language and how humans process phonemes. Spoken language isn't quite unambiguously causal so there's a limit to how far you can go for a given accuracy. I don't know where the efficiency curve is though. It wouldn't surprise me if 100ms was pushing it.
- moffkalast 7mo agoYeah the metric would be the total processing latency after that. I've found that VAD is honestly harder to get right than STT and if that fails, STT only gets garbage to process. Even humans sometimes have issues figuring out when exactly someone is done talking.
- jasonni 7mo agoThe python MLX version of Parakeet indeed support streaming: https://github.com/senstella/parakeet-mlx https://github.com/senstella/parakeet-mlx It requires modification of the inference algorithm. In this implementation, I see the author even uses a custom metal kernerl to get maximum performance. The Parakeet model batch inference logic is simple. But for streaming, it may require some effort to get the best performance. It's not only the depencency issue.
- ahaferburg 7mo agoAgree about the latency requirement. There's https://kyutai.org/stt https://kyutai.org/stt, which is very low latency. But it seems not as hackable.
- MarcLore 7mo ago[dead]
- pzo 7mo agoYou probably still better use inference on ANE (Apple Neural Engine) via CoreML rather than Metal - speed will be either similar or even faster on non-pro macbooks or iphones and power consumption significantly better. Metal or even MLX format doesn't have to be the fastest and the only way to access ANE is via CoreML. Can use this library: https://github.com/FluidInference/FluidAudio https://github.com/FluidInference/FluidAudio
- noahkay13 7mo agoThe CoreML backend is WIP in Axiom and will roll over to parakeet.cpp when it's ready, the same with CUDA. FluidAudio is a great option for those building Mac-only apps, but the goal with Axiom and Parakeet.cpp is to be very portable and embeddable into almost any app. I will write C and Swift wrappers shortly, then if it's really wanted, a Python wrapper.
- d4rkp4ttern 7mo agoFor MacOS I haven’t seen any STT app that has faster transcription than Hex (with Parakeet V3), which leverages Apple silicon + FluidAudio: https://github.com/kitlangton/Hex https://github.com/kitlangton/Hex This is now my standard way to speak to coding agents. I used to use Handy but Hex is even faster. Last I checked, Handy has stuttering issues but Hex doesn’t.
- mpalmer 7mo agoNot sure what the stuttering issues are... For my part, I don't need an app that's faster than Handy, and I do like that Handy is Tauri (Rust + web), which means it could be fully cross-platform eventually. It's mostly that the stack is just more hackable for me personally.