4 ms·
In my experience, Whisper does a great job even with specialized terminology. It won't catch everything, but I think it will exceed your expectations. One of th
by coder543 4y ago
In my experience, Whisper does a great job even with specialized terminology. It won't catch everything, but I think it will exceed your expectations. One of the hardest things about Whisper is choosing which model to use; they offer a variety of sizes, and sometimes the smaller ones do better than the larger ones. It's worth trying a few different models and deciding what is best for each particular application.
I will also say that I've personally been unimpressed with the new "large-v2" model, even though it supposedly scores better. The original "large-v1" model seems to work better than the "large-v2" model in the audio clips I've been testing Whisper against, but results will vary. In general, I find I'm really happy with what "small.en" and "medium.en" will emit, and they're much faster than the large models. (The ".en" models are specialized to English, and usually perform better for strictly English input, whereas the non-".en" models are trained on multiple languages.)
- simonw 4y agoI'd still like the ability to prime Whisper. I used it to transcribe a podcast episode I appeared on recently and one of the fixes I had to make was that ChatGPT came out as "chat GPT" every time it was mentioned: https://simonwillison.net/2023/Mar/7/kqed-forum/#kqed-forum https://simonwillison.net/2023/Mar/7/kqed-forum/#kqed-forum Update: turns out this exists already: https://platform.openai.com/docs/guides/speech-to-text/prompting https://platform.openai.com/docs/guides/speech-to-text/promp...
- coder543 4y agoI recorded myself saying a few sentences from that transcript, then fed it through different Whisper models. "small.en" and "large-v1" both generated "chat GPT", "large-v2" generated "chat-gpt", but somehow "medium.en" correctly generated "ChatGPT". This was the same audio sample fed through each of those four models, with no "prompting" as you're discussing. If I add "--initial_prompt ChatGPT", then all four models are able to get the spelling correct. Regardless, I don't think "chat GPT" versus "ChatGPT" is a huge deal. There will always be some level of uncertainty and ambiguity in the transcript, and even books written by humans always have a few typos get past multiple stages of copy editing. Perfection is virtually unachievable, but you can always scroll through the transcript and make some edits after the fact, if desired. Maybe some future model will magically eliminate all typos.
- simonw 4y agoYeah it wasn't a big problem for me - I had to do a bunch of other tidy-ups on the transcript anyway to add things like the name of the person who was speaking. I cleaned that bit up with a bulk replace of "chat GPT" with "ChatGPT" in VS Code.