8 ms·
Faster Whisper Transcription with CTranslate2
- rapsey 3y agoHopefully someone also makes it useful and ports it to a c/c++ lib like whisper.cpp
- Rastonbury 3y agoWhy is it useless if not?
- rapsey 3y agoIt is only useful if you are shipping a python app. But you are not doing that on mobile or the vast majority of desktop apps.
- barrenko 3y agoCan someone try to explain this architecture / whisper reimplementation? Thanks.
- guillaumekln 3y agoThe original Whisper implementation from OpenAI uses the PyTorch deep learning framework. On the other hand, faster-whisper is implemented using CTranslate2 [1] which is a custom inference engine for Transformer models. So basically it is running the same model but using another backend, which is specifically optimized for inference workloads. [1] https://github.com/OpenNMT/CTranslate2 https://github.com/OpenNMT/CTranslate2
- barrenko 3y agoMuch appreciated.
- aargh_aargh 3y agoThe project page mentions whisper-diarization (speaker recognition) as a user of faster-whisper. I've been in the market for that, definitely going to try it out. https://github.com/MahmoudAshraf97/whisper-diarization https://github.com/MahmoudAshraf97/whisper-diarization
- garblegarble 3y agoThere's also https://github.com/akashmjn/tinydiarize https://github.com/akashmjn/tinydiarize which received support recently in core whisper.cpp - it doesn't support the large model yet, but the results are promising
- 0cf8612b2e1e 3y agoThat’s a pretty exciting development. Is there any putative timelines on when it might get expanded to the large models? At this point, I would have to do a two pass workflow: first with large for accuracy and a second with the tube to assign speakers.
- garblegarble 3y agoNot that I can see, the developer's roadmap[1] currently is at writing a blogpost about it & trying different sampling methods, expanding to the large model looks to be a long way off (which is a real pity, even if there was just a finetune of large for english that would be a big help over the existing small english finetune). Could you go into more detail about your workflow? I'd been considering a two-pass approach myself until I discovered tinydiarize mentioned in whisper.cpp's --help text 1: https://github.com/akashmjn/tinydiarize#roadmap https://github.com/akashmjn/tinydiarize#roadmap
- 0cf8612b2e1e 3y agoI have not yet done this, but my thinking is that the diarization process would be useful data, but not so much that I would be willing to accept the quality dip of using the tiny trained model. I already run the large model at the word level, so I have high quality transcription + timestamps. Then I would perform the second whisper-pass with the diarization and discard all of the output except when a speaker change is identified. Merge the two results based on timestamps and there you go.
- nullandvoid 3y agoAte there any plans / usage examples of using this behind a Web app (node/react is my go to stack at the minute). I'm currently using openai whisper rest API, but sporadic latency spikes are a problem.
- aidenn0 3y agoHas someone already compared to whisper.cpp running on the CPU with cuBLAS on the GPU? That was a lot more than 4x faster than whisper on the CPU, and was sufficiently parsimonious with GPU ram to allow me running even the large model (and without quitting firefox, which is a bit of a GPU ram hog).