3 ms·
On Windows, I use scoop.sh: https://scoop.sh/#/apps?q=whisper https://scoop.sh/#/apps?q=whisper I was able to do this: scoop install main/whisper-cpp
by Leftium 2y ago
On Windows, I use scoop.sh: https://scoop.sh/#/apps?q=whisper https://scoop.sh/#/apps?q=whisper
I was able to do this:
scoop install main/whisper-cpp
mkdir models
## Download model file ggml-base.en.bin to models directory above
yt-dlp.exe -x --audio-format wav --audio-quality 16K -o "out.wav" ZMklf0vUl18
# Wrangle wav into 16kHz format (param above did not seem to work...)
ffmpeg -i out.wav -ar 16000 out-16kHz.wav
whisper.exe out-16kHz.wav
- llimllib 2y agoPresumably that used whisper's bundled tiny model, which is no better than youtube CC. A beef I have with whisper-cpp is that they totally refuse to handle model management. With mlx_whisper, I just have to tell it to use a model and it will download it if it's not already present: https://github.com/llimllib/yt-transcribe/blob/244841f83d83304710d1cedba5484cc53de61c9b/yt-transcribe#L236 https://github.com/llimllib/yt-transcribe/blob/244841f83d833... so if I add whisper.cpp as a dependency, I also have to add huggingface-cli or something similar. It also seems like huggingface-cli is not available on scoop
- Leftium 2y agoYes, the model must be downloaded separately (see my edited comment with bash commands/comments). The model is specified via whisper.exe `--model FNAME` parameter. By default, it looks for `models/ggml-base.en.bin`, but even that model must be downloaded separately. So you could do this: # Assumes ggml-large-v3.bin model file[1] was already downloaded to models/ folder whisper.exe --model models/ggml-large-v3.bin out-16kHz.wav [1]: https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-large-v3.bin https://huggingface.co/ggerganov/whisper.cpp/blob/main/ggml-...
- Leftium 2y agoNot as convenient, but you could also have the user manually install the model, like whisper does. Just forward the error message output by whisper, or even make a more user-friendly error message with instructions on how/where to download the models. Whisper does provide a simple bash script to download models: https://github.com/ggerganov/whisper.cpp/blob/master/models/download-ggml-model.sh https://github.com/ggerganov/whisper.cpp/blob/master/models/... (As a Windows user, I can run bash scripts via Git Bash for Windows[1]) [1]: https://git-scm.com/download/win https://git-scm.com/download/win
- llimllib 2y agothanks for all the help, I appreciate it.
- Leftium 2y agoWell, thanks to you I found out whisper generates decent audio transcriptions using a local LLM (relatively) easily, even on my 6+ year-old laptop. (I used to upload videos to YouTube just to get the auto captions.) I did some investigation, and it would not be difficult to convert the whisper LRC subtitle output into the format my fork of oTranscribe expects. I already made a simple tool to convert YouTube TTML/SBV subtitle output: https://github.com/Leftium/otrgen https://github.com/Leftium/otrgen
- llimllib 2y agothat's great! whisper is awesome software. I'm working on a golang version that links to whisper.cpp directly to maybe make porting easier/possible