8 ms·
Show HN: Gemini LLM corrects ASR YouTube transcripts
- alsetmusic 2y agoSeems like one of the places where LLMs make a lot of sense. I see some boneheaded transcriptions in videos pretty regularly. Comparing them against "more-likely" words or phrases seems like an ideal use case.
- petesergeant 2y agoAlso useful I think for checking human-entered transcriptions, which even on expensively produced shows, can often be garbage or just wrong. One human + two separate LLMs, and something to tie-break, and we could possibly finally get decent subtitles for stuff.
- leetharris 2y agoA few problems with this approach: 1. It brings everything back to the "average." Any outliers get discarded. For example, someone who is a circus performer plays fetch with their frog. An LLM would think this is an obvious error and correct it to "dog." 2. LLMs want to format everything as internet text which does not align well to natural human speech. 3. Hallucinations still happen at scale, regardless of model quality. We've done a lot of experiments on this at Rev and it's still useful for the right scenario, but not as reliable as you may think.
- falcor84 2y agoRegarding the frog, I would assume that the way to address this would be to feed the LLM screenshots from the video, if the budget allows.
- leetharris 2y agoGenerally yes. That being said, sometimes multimodal LLMs show decreased performance with extra modalities. The extra dimensions of analysis cause increased hallucination at times. So maybe it solves the frog problem, but now it's hallucinating in another section because it got confused by another frame's tokens. One thing we've wanted to explore lately has been video based diarization. If I have a video to accompany some audio, can I help with cross talk and sound separation by matching lips with audio and assign the correct speaker more accurately? There's likely something there.
- orion138 2y agoGoogle published Looking to Listen a while back. https://research.google/blog/looking-to-listen-audio-visual-speech-separation/ https://research.google/blog/looking-to-listen-audio-visual-...
- ldenoue 2y agoDo you have something to read about your study, experiments? Genuinely interested. Perhaps the prompts can be made to tell the LLM it's specifically handling human speech, not written text?
- devmor 2y agoThose transcriptions are already done by LLMs in the first place - in fact, audio transcription was one of the very first large scale commercial uses of the technology in its current iteration. This is just like playing a game of markov telephone where the step in OP's solution is likely higher compute cost than the step YT uses, because YT is interested in minimizing costs.
- albertzeyer 2y agoProbably just "regular" LMs, not large LMs, I assume. I assume some LM with 10-100M params or so, which is cheap to use (and very standard for ASR).
- devmor 2y agoCould be. I ran through some offline LMs for voice assisted home automation a couple years ago and they were subpar compared to even the pathetic offering that Youtube provides - but Google of course has much more focused resources to fine tune a small dataset model.
- dylan604 2y agoWhat about the cases where the human speaking is actually using nonsense words during a meandering off topic bit of "weaving"? Replacing those nonsense words would be a disservice as it would totally change the tone of the speech.
- dr_dshiv 2y agoThe first time I used Gemini, I gave it a youtube link and asked for a transcript. It told me how I could transcribe it myself. Honestly, I haven't used it since. Was that unfair of me?
- robrenaud 2y agoGemini is much worse as a product than 4o or Claude. I recommend using it from Google AI studio rather than the official consumer facing interface. But for tasks with large audio/visual input, it's better than 4o or Claude. Whether you want to deal with it being annoying is your call.
- Spooky23 2y agoThe consumer Gemini is very prudish and optimized against risk to Google.
- andai 2y agoGPT told me the same thing when I asked it to make an API call, or do an image search, or download a transcript of a YouTube video, or...
- replwoacause 2y agoNo it’s a terrible product that is embarrassingly bad compared to the competition. I ditched it after paying for a month of Gemini Advanced because it was so much worse other offerings.
- jazzyjackson 2y agoThinking about that time Berkeley delisted thousands of recordings of course content as a result of a lawsuit complaining that they could not be utilized by deaf individuals. Can this be resolved with current technology? Google's auto captioning has been abysmal up to this point, I've often wondered what the cost would be for google to run modern tech over the entire backlog of youtube. At least then they might have a new source of training data. https://news.berkeley.edu/2017/02/24/faq-on-legacy-public-course-capture-content/ https://news.berkeley.edu/2017/02/24/faq-on-legacy-public-co... Discussed at the time (2017) https://news.ycombinator.com/item?id=13768856 https://news.ycombinator.com/item?id=13768856
- delusional 2y agoThat's a legal issue. If humans wanted that content to be up, we just could have agreed to keep it up. Legal issues don't get solved by technology.
- jazzyjackson 2y agoWell. The legal complaint was that transcripts don't exist. The issue was that it was prohibitively expensive to resolve the complaint. Now that transcription is 0.1% of the cost it was 8 years ago, maybe the complaint could have been resolved. Is building a ramp to meet ADA requirements not using technology to solve a legal issue?
- delusional 2y agoNowhere on the linked page at least does it say that it was due to cost. It would seem more likely to me that it was a question of nobody wanting to bother standing up for the videos. If nobody wants to take the fight, the default judgement becomes to take it down. Building a ramp solves a problem. Pointing at a ramp 5 blocks away 7 years later and asking "doesn't this solve this issue" doesn't.
- pests 2y agoYet this feels very harrison bergeron to me. To handicap those with ability so we all can be at the same level.
- wood_spirit 2y agoAs an aside, has anyone else had some big hallucinations with the Gemini meet summaries? Have been using it a week or so and loving the quality of the grammar of the summary etc, but noticed two recurring problems: omitting what was actually the most important point raised, and hallucinating things like “person x suggested y do z” when, really, that is absolutely the last thing x would really suggest!
- hunter2_ 2y agoIt can simultaneously be [the last thing x would suggest] and [a conclusion that an uninvolved person tasked with summarizing might mistakenly draw, with slightly higher probability of making this mistake than not making it] and theoretically an LLM attempts to output the latter. The same exact principle applies to missing the most important point.
- leetharris 2y agoThe Google ASR is one of the worst on the internet. We run benchmarks of the entire industry regularly and the only hyperscaler with a good ASR is Azure. They acquired Nuance for $20b a while ago and they have a solid lead in the cloud space. And to run it on a "free" product they probably use a very tiny, heavily quantized version of their already weak ASR. There's lots and lots of better meeting bots if you don't mind paying or have low usage that works for a free tier. At Rev we give away something like 300 minutes a month.
- baxtr 2y agoVery interesting. Thanks for sharing. Since you have experience in this, I’d like to hear your thoughts on a common assumption. It goes like this: don’t build anything that would be feature for a Hyperscalar because ultimately they win. I guess a lot of it is a question of timing?
- leetharris 2y agoI think it really depends on whether or not you can offer a competitive solution and what your end goals are. Do you want an indie hacker business, do you want a lifestyle business, do you want a big exit, do you want to go public, etc? It is hard to compete with these hyperscalers because they use pseudo anti-competitive tactics that honestly should be illegal. For example, I know some ASR providers have lost deals to GCP or AWS because those providers will basically throw in ASR for free if you sign up for X amount of EC2 or Y amount of S3, services that have absurd margins for the cloud providers. Still, stuff like Supabase, Twilio, etc show there is a market. But it's likely shrinking as consolidation continues, exits slow, and the DOJ turns a blind eye to all of this.
- leetharris 2y agoThe main challenge with using LLMs pretrained on internet text for transcript correction is that you reduce verbatimicity due to the nature of an LLM wanting to format every transcript as internet text. Talking has a lot of nuances to it. Just try to read a Donald Trump transcript. A professional author would never write a book's dialogue like that. Using a generic LLM on transcripts almost always reduces accuracy as a whole. We have endless benchmark data to demonstrate this at RevAI. It does, however, help with custom vocabulary, rare words, proper nouns, and some people prefer the "readability" of an LLM-formatted transcript. It will read more like a wikipedia page or a book as opposed to the true nature of a transcript, which can be ugly, messy, and hard to parse at times.
- phrotoma 2y agoI googled "verbatimicity" and all I could find was stuff published by rev.ai which didn't (at a quick glance) define the term. Can you clarify what this means?
- depr 2y agoMost likely they mean the degree of being verbatim or exact in reproduction.
- dylan604 2y ago> A professional author would never write a book's dialogue like that. That's a bit too far. Ever read Huck Finn?
- icelancer 2y agoNice use of an LLM - we use Groq 70b models for this in our pipelines at work. (After using WhisperX ASR on meeting files and such) One of the better reasons to use Cerebras/Groq that I've found so you can return huge amounts of clean text back fast for processing in other ways.
- ldenoue 2y agoAlthough Gemini accepts very long input context, I found that sending more than 512 or so words at a time to the LLM for "cleaning up the text" yields hallucinations. That's why I chunk the raw transcript into 512-word chunks. Are you saying it works with 70B models on Groq? Mixtral, Llama? Other?
- icelancer 2y agoYeah, I've had no issues sending tokens up to the context limit. I cut it off with a 10% buffer but that's just to ensure I don't run into tokenization miscounting between tiktoken and whatever tokenizer my actual LLM uses. I have had little success with Gemini and long videos. My pipeline is video -> ffmpeg strip audio -> whisperX ASR -> groq (L3-70b-specdec) -> gpt-4o/sonnet-3.5 for summarization. Works great.
- bob_theslob646 2y agoWhen you did this, I am assuming you cut the audio off around 5 mins? https://github.com/google-gemini/generative-ai-js/issues/269#issuecomment-2425235602 https://github.com/google-gemini/generative-ai-js/issues/269...
- tombh 2y agoASR: Automatic Speech Recognition
- joshdavham 2y agoI was too afraid to ask!
- throwaway106382 2y agoNot to be confused with "Autonomous Sensory Meridian Response" (ASMR) - a popular category of video on Youtube.
- hackernewds 2y agoHow would they be confused?
- xanth 2y agoThis was a clever jape; a good example of a ironic anti-humor. But I don't think you were confused by that ether ;)
- djmips 2y agoclever japes are not desired on HN - there's Reddit for that my friend.
- wodenokoto 2y agoI can't explain the how, but I thought it was the ASMR thing the title referred to.
- throwaway106382 2y agoI think more people actually know what ASMR is as opposed to ASR. Lots of ASMR videos are people speaking/whispering at extremely low volume. I don't think it's quite out of the realm of the possibility to have interpreted as "Gemini LLM corrects ASMR YouTube transcripts". Because you know..they're whispering so might be hard to understand or transcribe.
- sorenjan 2y agoUsing an LLM to correct text is a good idea, but the text transcript doesn't have information about how confident the speech to text conversion is. Whisper can output confidence for each word, this would probably make for a better pipeline. It would surprise me if Google doesn't do something like this soon, although maybe a good speech to text model is too computationally expensive for Youtube at the moment.
- dylan604 2y agoDepends on your purpose of the transcript. If you are expecting the exact form of the words spoken in written form, then any deviation from that is no longer a transcription. At that point it is text loosely based on the spoken content. Once you accept it okay for the LLM to just replace words in a transcript, you might as well just let it make up a story based on character names you've provided.
- falcor84 2y ago> any deviation from that is no longer a transcription That's a wild exaggeration. Professional transcripts often have small (and not so small) mistakes, caused by typos, mishearing or lack of familiarity with the subject matter. Depending on the case, these are then manually proofread, but even after proofreading, some mistakes often remain, and occasionally even introduced.
- dylan604 2y agomaybe, but typos are not even the same thing as an LLM thinking of better next choice in words than actually just transcribing what was heard.
- kelvinjps 2y agoGoogle should have the needed tech for good AI transcription, why the don't integrate them in their auto-captioning? and instead the offer those crappy auto subtitles
- briga 2y agoAre they crappy though? Most of the time it gets things right, even if they aren't as accurate as a human. And sure, they probably have better techniques for this, but are they cost-effective to run at YouTube-scale? I think their current solution is good enough for most purposes, even if it isn't perfect
- InsideOutSanta 2y agoI'm watching YouTube videos with subtitles for my wife, who doesn't speak English. For videos on basic topics where people speak clear, unaccented English, they work fine (i.e. you usually get what people are saying). If the topic is in any way unusual, the recording quality is poor, or people have accents, the results very quickly turn into a garbled mess that is incomprehensible at best, and misleading (i.e. the subtitles seem coherent, but are wrong) at worst.
- wahnfrieden 2y agoJapanese auto captions suck
- summerlight 2y agoYT is using USM, which is supposed to be their SOTA ASR model. Gemini have much better linguistic knowledge, but it's likely prohibitively expensive to be used on all YT videos uploaded everyday. But this "correction" approach seems to be a nice cost-effective methodology to apply LLM indeed.
- Timwi 2y agoCan I use this to generate subtitles for my own videos? I would love to have subtitles on them but I can't be bothered to do all the timing synchronization by hand. Surely there must be a way to automate that?
- geor9e 2y agoThat's called Youtube Automatic Speech Recognition (captioning), and is what this tool uses as input. You can turn those on in youtube studio.
- sidcool 2y agoThis is pretty cool. But at the risk of a digression, I can't imagine sharing my API keys with a random website on HN. There has to be a safe approach to this. Like limited use API keys, rate limited API keys or unsafe API keys etc.
- thomasahle 2y agoCan't you just create a new API key with a limited budget?
- ldenoue 2y agoI should do that, let me try.
- sidcool 2y agoThe risk of leakage is very high. If Anthropic, Google, OpenAI can provide dispensible keys, it will be great.
- thomasahle 2y agoBoth OpenAI and Anthropic let you disable and delete keys. I'd be surprised if Google doesn't.
- mst 2y agoI'm aware this isn't a *proper* solution, but "throw your current API key at it, then as soon as you're done playing around, execute a test of your API key rotation scripting" isn't a terrible workaround, especially if you're the sort of person who really *meant* to have tested said scripting recently but kept not getting around to it ("hi").
- pachico 2y agoHmm, so this is expecting me to upload a personal API Key...
- ldenoue 2y agoIt’s not uploaded anywhere: the client fetches Gemini servers directly from your browser. But I understand it can be difficult to trust: that’s why the project is on GitHub so you can run it on your own machine and look at how the key is used. I will try to offer a version that doesn’t require any key.
- replwoacause 2y agoIn my experience Gemini Advanced is still so far behind ChatGPT and Claude. Recently it flat out refused to answer my fairly straightforward question by saying “I am just a large language model and cannot help you with that”. The conversation was totally benign but it flat out shit the bed so I canceled my subscription right then and there.
- ldenoue 2y agoDid you see this using the API or the online product gemini?
- replwoacause 2y agoThe online product, I haven’t tried the API.