5 ms·
Exciting to see some parts now (optionally) running on the GPU.
by simse 3y ago
Exciting to see some parts now (optionally) running on the GPU.
- rektide 3y agoFull circle eh. I wonder how well it compares to just trying to use the actual Whisper models on a variety of existing GPU capable bigger frameworks. I don't know much practically about how hard it would be to take the Whisper PyTorch (1 or 2?) trained models & to make good use of them elsewhere. I expect Whisper.cpp probably better caters to users, is more readily consumable. Doing the same integer quantization that 1.4.0 whisper.cpp does would be another ask: how hard would that be? Fwiw, Whisper.cpp uses Nvidia's cuBLAS. There does appear to be an AMD rocm port. https://github.com/ROCmSoftwarePlatform/rocBLAS https://github.com/ROCmSoftwarePlatform/rocBLAS
- angrais 3y ago>> I don't know much practically about how hard it would be to take the Whisper PyTorch (1 or 2?) trained models & to make good use of them elsewhere. What do you mean by "elsewhere"? The whisper models can be used without much effort. The tooling provided by OpenAI allows for them to be used and installed in python with ease. People use them in companies to transcribe internal media datasets. No need for CPP implementation for internal uses. Main problem is when _on device_ is required and the OpenAI model is not quantised and so use on say a smartphone becomes trickier.