3 ms·
Gemini 1.5 Pro multimodal with audio file input plus a text prompt that asks for a transcription. You can add all the relevant context you know to the text prom
by mbrock 2y ago
Gemini 1.5 Pro multimodal with audio file input plus a text prompt that asks for a transcription. You can add all the relevant context you know to the text prompt. You can also chat with it before asking for the full transcription and clear up any confusions. There are no products for this yet, it's a new capability.
I just started trying this today with my wife's hour long interviews in Latvian language and it is extremely good, far better than any transcription model. This is a huge SoTA LLM with audio tokens, so it just has vastly more capability than Whisper or whatever. In my case it nails all kinds of brand names, weird neologisms and loan words, it writes inline quotation marks when the speaker is quoting someone, and so on.
This is what GPT-4o supposedly can do in the version OpenAI has postponed rolling out.
If you want, I can try to do it for you if you send the audio.
- mbrock 2y agoI've noticed it's necessary to cut long audio into sections, otherwise it starts getting confused and repeating itself. The output token limit only lets you transcribe a few pages at a time anyway.