3 ms·
This was a breeze to install on Linux. However, I haven't managed to get realtime transcription working yet, ala Whisper.cpp stream or Moonshine. --from-mic on
by Curiositry 8mo ago
This was a breeze to install on Linux. However, I haven't managed to get realtime transcription working yet, ala Whisper.cpp stream or Moonshine.
--from-mic only supports Mac. I'm able to capture audio with ffmpeg, but adapting the ffmpeg example to use mic capture hasn't worked yet:
ffmpeg -f pulse -channels 1 -i 1 -f s16le - 2>/dev/null | ./voxtral -d voxtral-model --stdin
It's possible my system is simply under spec for the default model.
I'd like to be able to use this with the voxtral-q4.gguf quantized model from here: https://huggingface.co/TrevorJS/voxtral-mini-realtime-gguf https://huggingface.co/TrevorJS/voxtral-mini-realtime-gguf
- yjftsjthsd-h 8mo agoDoes it work if you use ffmpeg to feed it audio from a file? I personally would try file->ffmpeg->voxtral then mic->ffmpeg->file, and then try to glue together mic->ffmpeg->voxtral. (But take with grain of salt; I haven't tried yet)
- Curiositry 8mo agoRecording audio with FFMPEG, and transcribing a file that’s piped from FFMPEG both work. Given that it took 19.64 mins to transcribe the 11 second sample wav, it’s possible I just didn’t wait long enough :)
- yjftsjthsd-h 8mo agoAh. In that case... Yeah. Is it using GPU, and does the whole model fit in your (V)RAM?
- ekianjo 8mo agoThis is a CPU implementation only.
- yjftsjthsd-h 8mo agoOh, that's interesting. The readme talks about GPU acceleration on Apple Silicon and I didn't see anything explicit for other platforms, so I assumed it needs GPU everywhere, but it does BLAS acceleration which a web search seems to agree is just a CPU optimized math library. That's great; should really increase the places where it's useful:)
- ekianjo 8mo agoIt should be possible to develop a cuBLAS backend to accelerate BLAS on Nvidia.
- jwrallie 8mo agoI am interested in a way to capture audio not only from the mic, but also from one of the monitor ports so you could pipe the audio you are hearing from the web directly for real-time transcription with one of these solutions. Did anyone manage to do that? I can, for example, capture audio from that with Audacity or OBS Studio and do it later, so it should be possible to do it in real time too assuming my machine can keep up.
- bebna 8mo agoSet -i 1 to -i default or to one of your monitors, look them up with pactl list short sources https://trac.ffmpeg.org/wiki/Capture/PulseAudio https://trac.ffmpeg.org/wiki/Capture/PulseAudio
- jandrese 8mo agoFrom my testing on Linux this model is way too slow for anything close to realtime. The machine I’m using is kinda old, but a 12 minute input file took half a day to process.