4 ms·
Ask HN: Which speech to text model would you recommend?
I may need to perform a bit of speech-to-text (English at least, but in perspective - multilingual also) from video or audio files.
Which speech-to-text model/API would you recommend,
which sort of performs the best and can also do noise etc reduction?
- smoldesu 3y agoWhisper, 100%. It's small, fast and does a really good job with most of the recordings I can feed it. IIRC, there are both English and mixed-language models to choose from as well.
- spacetimeuser5 3y agoThis one https://huggingface.co/openai/whisper-large-v3 https://huggingface.co/openai/whisper-large-v3 ? Do I need (and where from) to use CUDA from some google.colab or 4-core AMD Ryzen 3 4300U CPU on a laptop will handle it for an initial test program?
- smoldesu 3y agoThat one should work. The model itself is extremely fast without any acceleration, in my experience. The AMD CPU should perform just fine. Good luck!
- spacetimeuser5 3y agoIs AWS speech-to-text (free tier) worth it?