4 ms·
> but Google seems to be pretty honest here. They're not. Take "This means that AI Overviews generally don't “hallucinate” or make things up in the ways that
by ADeerAppeared 2y ago
> but Google seems to be pretty honest here.
They're not.
Take "This means that AI Overviews generally don't “hallucinate” or make things up in the ways that other LLM products might."
This is just a lie, in two ways:
1. It's still an LLM, feeding search results into the prompts and asking for a summary reduces the probability of hallucination compared to directly asking the question in the prompt, it's still a probability non-negligibly above zero. This kind of pretending you can "fix" the hallucination problem by feeding data into the prompt is extremely wrong and dangerous.
2. It's missing the entire point.
"Um akshually the LLM didn't hallucinate, we simply gave it Reddit user Fucksmith's post which it confidently restated as truth" (or indeed, the much funnier, "we gave it The Onion, which it confidently restated as truth") is functionally the same as a hallucination to the end user. The mechanism here does not matter.
These are both fundamental critical issues to AI overviews. "Rare" hallucinations are unacceptable on a tool used billions of times a day. And the LLM paraphrasing lies/satire/ignorance/etc into truth is a fundamental flaw of using LLMs to paraphrase like this.
You can't just patch this issue, the entire thing is fundamentally flawed.
- Ukv 2y ago> This is just a lie [...] it's still a probability non-negligibly above zero. This kind of pretending you can "fix" the hallucination problem [...] I don't think "generally don't hallucinate" and "it's usually for other reasons" are implying a total fix. From what I've seen, it is true that these errors stemmed from uncritically accepting satire/unreliable sources rather than hallucinations. > It's missing the entire point. "Um akshually the LLM didn't hallucinate, we simply gave it Reddit user Fucksmith's post which it confidently restated as truth" (or indeed, the much funnier, "we gave it The Onion, which it confidently restated as truth") is functionally the same as a hallucination to the end user. The mechanism here does not matter. When you can check the provided source it's summarising and see that it's a satire website, I think that's materially different than ChatGPT's hallucinations from thin air. It is still an issue, and if the post was just deflecting by saying "it's not typical hallucination so it doesn't matter" then I'd agree that'd be insufficient, but it seems totally fine to clarify that it's a different issue and go on to describe how they're working on fixing that issue. > "Rare" hallucinations are unacceptable on a tool used billions of times a day I don't think 100.0% accuracy is a reasonable bar for anything, including humans. My benchmark would be roughly "people come away with the correct information the same or more often as they did before, with features like highlighted snippets".
- ADeerAppeared 2y ago> When you can check the provided source it's summarising and see that it's a satire website, I think that's materially different than ChatGPT's hallucinations from thin air. I disagree. The moment you accept the need to always check the source, there's no fucking point to this anymore. If you have to check the source ... just provide the links like Google Search does. There's no point in a summary whose full original text you always have to read. > From what I've seen, it is true that these errors stemmed from uncritically accepting satire/unreliable sources rather than hallucinations. And so, this is why I don't think this holds up; If we consider this a failure mode of the user, the entire tool is pointless. It's plainly obvious that the intended benefit of AI Overview is not having to read the pages of the search result. And with that intent, this is not a failure of the user.
- Ukv 2y ago> The moment you accept the need to always check the source, there's no fucking point to this anymore. > If you have to check the source ... just provide the links like Google Search does. There's no point in a summary whose full original text you always have to read. > [...] It's plainly obvious that the intended benefit of AI Overview is not having to read the pages of the search result. I see its purpose as similar to the text snippets and "highlight" box that Google search has had for a while: quickly picking out information relevant to your query. I don't think most answers actually warrant the user to verify that it came from a reliable source. If you ask "what foods end with um" or "good ideas for birthday party" you can generally just judge the answers yourself, like you would even if you already knew the answer came from Quora/reddit. But for answers that do (which is still plenty), it makes it easier to click through and check the part of the summary you determined to be relevant, as opposed to manually parsing through pages to find that part in the first place. > And so, this is why I don't think this holds up; If we consider this a failure mode of the user, the entire tool is pointless. I don't consider it solely a failure on the user's part - the tool is failing here and Google claim to be making changes to improve it. I don't think lack of 100% accuracy makes automatic summaries pointless - they are still useful for surfacing information.