5 ms·
Pure C, CPU-only inference with Mistral Voxtral Realtime 4B speech to text model
- Curiositry 8mo agoThis was a breeze to install on Linux. However, I haven't managed to get realtime transcription working yet, ala Whisper.cpp stream or Moonshine. --from-mic only supports Mac. I'm able to capture audio with ffmpeg, but adapting the ffmpeg example to use mic capture hasn't worked yet: ffmpeg -f pulse -channels 1 -i 1 -f s16le - 2>/dev/null | ./voxtral -d voxtral-model --stdin It's possible my system is simply under spec for the default model. I'd like to be able to use this with the voxtral-q4.gguf quantized model from here: https://huggingface.co/TrevorJS/voxtral-mini-realtime-gguf https://huggingface.co/TrevorJS/voxtral-mini-realtime-gguf
- yjftsjthsd-h 8mo agoDoes it work if you use ffmpeg to feed it audio from a file? I personally would try file->ffmpeg->voxtral then mic->ffmpeg->file, and then try to glue together mic->ffmpeg->voxtral. (But take with grain of salt; I haven't tried yet)
- Curiositry 8mo agoRecording audio with FFMPEG, and transcribing a file that’s piped from FFMPEG both work. Given that it took 19.64 mins to transcribe the 11 second sample wav, it’s possible I just didn’t wait long enough :)
- yjftsjthsd-h 8mo agoAh. In that case... Yeah. Is it using GPU, and does the whole model fit in your (V)RAM?
- ekianjo 8mo agoThis is a CPU implementation only.
- yjftsjthsd-h 8mo agoOh, that's interesting. The readme talks about GPU acceleration on Apple Silicon and I didn't see anything explicit for other platforms, so I assumed it needs GPU everywhere, but it does BLAS acceleration which a web search seems to agree is just a CPU optimized math library. That's great; should really increase the places where it's useful:)
- ekianjo 8mo agoIt should be possible to develop a cuBLAS backend to accelerate BLAS on Nvidia.
- jwrallie 8mo agoI am interested in a way to capture audio not only from the mic, but also from one of the monitor ports so you could pipe the audio you are hearing from the web directly for real-time transcription with one of these solutions. Did anyone manage to do that? I can, for example, capture audio from that with Audacity or OBS Studio and do it later, so it should be possible to do it in real time too assuming my machine can keep up.
- bebna 8mo agoSet -i 1 to -i default or to one of your monitors, look them up with pactl list short sources https://trac.ffmpeg.org/wiki/Capture/PulseAudio https://trac.ffmpeg.org/wiki/Capture/PulseAudio
- jandrese 8mo agoFrom my testing on Linux this model is way too slow for anything close to realtime. The machine I’m using is kinda old, but a 12 minute input file took half a day to process.
- deleted 8mo ago[deleted]
- sgt 8mo agoI'm very interested in speech to text - but like tricky dialects and use of various terminologies but I'm still confused as to where to start in the best possible place, in order to train the models with a huge database of voice samples I own. Any ideas from the HN crowd currently involved in speech 2 text models?
- written-beyond 8mo agoFunny, this and the Rust runtime implementation are neck and neck on the frontpage right now. Cool project!
- genie3io 8mo ago[dead]
- hrpnk 8mo agoThere is also a MLX implementation: https://github.com/awni/voxmlx https://github.com/awni/voxmlx
- MORPHOICES 8mo ago[dead]
- mythz 8mo agoBig fan of Salvatore's voxtral.c and flux2.c projects - hope they continue to get optimized as it'd be great to have lean options without external deps. Unfortunately it's currently too slow for real-world use (AMD 7800X3D/Blas) when adding Voice Input support to llms-py [1]. In the end Omarchy's new support for voxtype.io provided the nicest UX, followed by Whisper.cpp, and despite being slower, OpenAI's Whisper is still a solid local transcription option. Also very impressed with both the performance and price of Mistral's new Voxtral Transcription API [2] - really fast/instant and really cheap ($0.003/min), IMO best option in CPU/disk-constrained environments. [1] https://llmspy.org/docs/features/voice-input https://llmspy.org/docs/features/voice-input [2] https://docs.mistral.ai/models/voxtral-mini-transcribe-26-02 https://docs.mistral.ai/models/voxtral-mini-transcribe-26-02
- mijoharas 8mo agoOne thing I keep looking for is transcribing while I'm talking. I feel like I need that visual feedback. Does voxtype support that? (I wasn't able to find anything at glance) Handy claims to have an overlay, but it seems to not work on my system.
- mythz 8mo agoNot sure how it works in other OS's but in Omarchy [1] you hold down `Super + Ctrl + X` to start recording and release it to stop, while it's recording you'll see a red voice recording icon in the top bar so it's clear when its recording. Although as llms-py is a local web App I had to build my own visual indicator [2] which also displays a red microphone next to the prompt when it's recording. It also supports both Tap On/Off and hold down for recording modes. When using voxtype I'm just using the tool for transcription (i.e. not Omarchy OS-wide dictation feature) like: $ voxtype transcribe /path/to/audio.wav If you're interested the Python source code to support multiple voice transcription backends is at: [3] [1] https://learn.omacom.io/2/the-omarchy-manual/107/ai https://learn.omacom.io/2/the-omarchy-manual/107/ai [2] https://llmspy.org/docs/features/voice-input https://llmspy.org/docs/features/voice-input [3] https://github.com/ServiceStack/llms/blob/main/llms/extensions/voice/__init__.py https://github.com/ServiceStack/llms/blob/main/llms/extensio...
- sylware 8mo agoFinally a plain and simple C lib to run LLM opened weights?
- alextray812 8mo agoFrom a cybersecurity perspective, this project is impressive not just for performance, but for transparency.
- d4rkp4ttern 8mo agoI use the open source Handy [1] app with Parakeet V3 for STT when talking to coding agents and I’ve yet to see anything that beats this setup in terms of speed/accuracy. I get near instant transcription, and the slight accuracy drop is immaterial when talking to AIs that can “read between the lines”. I tried incorporating this Voxtral C implementation into Handy but got very slow transcriptions on my M1 Max MacBook 64GB. [1] https://github.com/cjpais/Handy https://github.com/cjpais/Handy I’ll have to try the other implementations mentioned here.
- thethimble 8mo agoHandy is great but I wish the STT was realtime instead of batch
- d4rkp4ttern 8mo agoThere’s a tradeoff here. If you want streaming output, then you lose the opportunity to clean it up in post processing such as removing filler words or removing stutters, etc., or any other AI based cleanup. The MacOS built-in dictation streams in real time and also does some cleanup, but it does awkward things, like the streaming text shows up at the bottom of the screen. Also I don’t think it’s as accurate as Parakeet V3, and there’s a start up lag of 1-2 secs after hitting the dictation shortcut, which kills it for me.
- thethimble 8mo agoI feel like this is a solvable problem. If you emit an errant word that should be replaced, why not correspondingly emit backspaces to just rewrite the word? I feel like this is the best of both worlds. Perhaps a little janky with backspaces, but still technically feasible.
- t0md4n 8mo agoHave you tried Hex? https://github.com/kitlangton/Hex https://github.com/kitlangton/Hex Faster than handy and uses way less memory.
- 9999_points 8mo agoIt seems so bizarre that we need a nearly 9gb model to do something you could do over 20 years ago with ~200mb.
- ks2048 8mo agoShould this work on a 16GB M3 MacBook Pro? It starts to load, but hangs or is too slow.
- BugsJustFindMe 8mo agoThe title here says CPU only, but that's wrong. The repo clearly says it has GPU acceleration and doesn't make any claims about CPUness.