5 ms·
What is the current SOTA for voice->text? I have a recording I've been sitting on for 2 years(a guest lecture which a friend recorded) which contains a very he
by IncreasePosts 2y ago
What is the current SOTA for voice->text?
I have a recording I've been sitting on for 2 years(a guest lecture which a friend recorded) which contains a very heavy amount of background noise, where you can just barely make out what is being said by the lecturer. I wonder if there is any hope I will ever be able to read a transcript from it.
I can figure out what the lecturer is saying (maybe only because I have some context about what he is talking about), but it is too painful to sit through 2 hours of it and try to transcribe it.
I tried uploading the audio file to this service, but basically get nothing useful returned to me.
- bequanna 2y agoI’ve had good results with Whisper. You can use the OpenAI API and it is also open source: https://github.com/openai/whisper https://github.com/openai/whisper
- geerlingguy 2y agoAnd if you have a Mac, get MacWhisper. It's been a godsend for transcribing almost anything. Usually pretty good if the main voice is discernible at all—though in OP's case, if the main voice is almost indistinguishable it might not do amazing.
- MengerSponge 2y agoI've had good luck with Buzz: https://github.com/chidiwilliams/buzz https://github.com/chidiwilliams/buzz
- pants2 2y agoDeepgram Nova 2 is among the best right now, more accurate than Whisper in my testing.
- xan_ps007 2y agoWe have built a dockerized open source stack for Whisper + Llama3 + MeloTTS. Whisper and MeloTTS for now works fine for our use cases. https://github.com/bolna-ai/bolna/tree/master/examples/whisper-melo-llama3 https://github.com/bolna-ai/bolna/tree/master/examples/whisp...
- Prompter9856 2y agoYou can upload your file and try out Deepgram for free, just to see what the results look like for your audio. No harm in trying: https://deepgram.com/free-transcription https://deepgram.com/free-transcription Disclosure: I work for Deepgram
- ProfessorLayton 2y agoTry giving Audacity a shot to cleanup the audio, it has a built-in noise reduction feature that's configurable. I've used it to varying degrees of success, but works especially well with the same sounds ANC headphones are good at blocking.
- duped 2y agofwiw, the commercial services that do this are called "audio forensics" (unsurprisingly, they're usually hired by cops and lawyers). You pay them to use their (often expensive) software tools to clean up audio and provide a transcription. I get the appeal of automating this task but the SOTA is not to automate it at all.
- willsmith72 2y agoyou can pay people online trivial amounts of money for this, it will be far cheaper and quicker than waiting for the right AI by the way, even once we get to a sufficient AI, how do you verify it without listening to the whole thing anyway? it's only 2 hours, if you're a fast typer at max it would take you 1 work day to transcribe yourself, or <$200 by a professional
- IncreasePosts 2y agoIt's not an issue of typing speed, it's that puzzling out what was said takes a number of re-listens, at high gain which hurts my ears after a while since the voice is just barely above the noise floor.
- dbspin 2y agoI don't know about SOTA, but 'Adobe Podcast Studio' a web app that I believe is still free / in beta offers excellent sound cleanup. So much so that many podcast / radio producers I know no longer frequently use Izotope RX - one of the industry standard tools. Adobe are obviously horrendous, but if its for a one time use I'd give it a go. The feature you want is the 'enhance speech filter'. https://podcast.adobe.com/enhance https://podcast.adobe.com/enhance
- wkcheng 2y agoThis is really helpful, thanks! I have a bunch of audio that I need to clean up and this looks like it could fit the bill. Do you know if there are any license issues with this? I don't see any license page--will they train/retain the recording?
- dbspin 2y agoI'm not sure - that's a really good question. I'd assume anything uploaded to a deep learning system retains the data for future training. But I have no information on the licensing of this tool specifically. Especially given the recent furore over Adobe's licensing terms.
- echelon 2y agoSTT: Whisper TTS: GPTSOVITS / StyleTTS2 VTV: RVCv2 Open source isn't really doing a great job at voice, music, or video. It's managing to keep up in LLM and image spaces, but it's falling far behind in the multimedia department.
- mbrock 2y agoGemini 1.5 Pro multimodal with audio file input plus a text prompt that asks for a transcription. You can add all the relevant context you know to the text prompt. You can also chat with it before asking for the full transcription and clear up any confusions. There are no products for this yet, it's a new capability. I just started trying this today with my wife's hour long interviews in Latvian language and it is extremely good, far better than any transcription model. This is a huge SoTA LLM with audio tokens, so it just has vastly more capability than Whisper or whatever. In my case it nails all kinds of brand names, weird neologisms and loan words, it writes inline quotation marks when the speaker is quoting someone, and so on. This is what GPT-4o supposedly can do in the version OpenAI has postponed rolling out. If you want, I can try to do it for you if you send the audio.
- mbrock 2y agoI've noticed it's necessary to cut long audio into sections, otherwise it starts getting confused and repeating itself. The output token limit only lets you transcribe a few pages at a time anyway.
- thatsadude 2y agoRecent ASR models are already robust to noise due to Spec augment and large-scale data. If you use these noise reduction services to remove noise, ASR models will have harder time to recognize denoised audio. The reason is that noise reduction will create distortion which ASR didn't see during training.
- sweetdreamerit 2y agoI would try, for 10 minutes, the following solution: you listen to the lecture and repeat every word the lecturer said. Then use the record of your voice for the voice 2 text process. If it works, do it for the whole talk.