4 ms·
I'm usually pretty cynical about these kinds of posts, but Google seems to be pretty honest here. It's a pretty reasonable take on what went wrong and what they
by programjames 2y ago
I'm usually pretty cynical about these kinds of posts, but Google seems to be pretty honest here. It's a pretty reasonable take on what went wrong and what they're working on to fix it. I still don't particularly like the AI Overview though.
- gmerc 2y agoExcuse me, lack of information is a misdirect. They have data warehouse of information that knows where this article is coming from, and every fucking url ever put into chrome. They bought the rights to /r/the_onion It’s an onion article. Directly from the page. Are they saying “the mighty google search” + “AI” is going to be somehow worse because they don’t have enough data about the fucking Onion article with dozens of hits in 2021? Even a 3.8B parameter edge model (phi3-mini) has enough satire token clustered around the first two paragraphs to flag it accordingly. https://cdn.some.pics/snekoil/66594bca93e19.jpg https://cdn.some.pics/snekoil/66594bca93e19.jpg as does llama3-7B https://cdn.some.pics/snekoil/66594cb8cddcd.jpg https://cdn.some.pics/snekoil/66594cb8cddcd.jpg as does GPT4 https://cdn.some.pics/snekoil/66594e1a3de29.jpg https://cdn.some.pics/snekoil/66594e1a3de29.jpg as does Claude https://cdn.some.pics/snekoil/665955fb22618.jpg https://cdn.some.pics/snekoil/665955fb22618.jpg Instead they gaslight us with words like “faithful” and “nobody asked the question”. Their entire competition executes better. Just because they AI wash the search data to boost their stock narrative we don’t have to accept their lame limitations narrative. They chose this fight. Also see my longer commentary at the root which is getting google brigaded like many other critical comments in this thread. Also, this is Zuck level of malevolent naïveté, it’s not like they don’t have experience https://www.theverge.com/2024/2/21/24079371/google-ai-gemini-generative-inaccurate-historical https://www.theverge.com/2024/2/21/24079371/google-ai-gemini...
- programjames 2y agoI think the difference between our takes is our understanding of how the "AI Overview" was trained. Google said, > AI Overviews work very differently than chatbots and other LLM products that people may have tried out and later, > Let’s take a look at an example: “How many rocks should I eat?” Prior to these screenshots going viral, practically no one asked Google that question. It seems it's being trained based on search queries. This is important to protect against confabulation, but obviously AI will fail on out-of-distribution data. Such as search queries that practically no one asked about before. Now, they could ask an LLM if the AI Overview seems reasonable, but the entire point of their alternate training scheme is to guarantee the output can be backed with sources. Why introduce a potential failure mechanism? You say, "they gaslight us with words like 'faithful' and 'nobody asked the question,'" but given this is how their model works, it seems like the best, non-gaslighty defense. Plus, it's pretty unreasonable to expect Google to fix every potential bug before production. As they said, > there’s nothing quite like having millions of people using the feature with many novel searches. We’ve also seen nonsensical new searches, seemingly aimed at producing erroneous results. Having "a content policy violation on less than one in every 7 million unique queries" seems pretty solid to me. I'm sure it'd be much higher if they were using an LLM like Claude. EDIT: My argument is pretty much, it looks like they designed this to be helpful and accurate, so the errors seem like honest mistakes from imperfect execution rather than deception or politicking. Their post seems pretty open about acknowledging what went wrong and that they want to do better.
- gmerc 2y agoSo they failed to make it work and are now trying to shift the goalposts. They should not have launched it then.
- prepend 2y ago> Having "a content policy violation on less than one in every 7 million unique queries" seems pretty solid to me. Not to me. AI will lead to more queries because I search differently with AI results than an index/page rank result. Also, it’s not about the unique queries, but the number of users running them. This seems like “data waving” where google is giving some not relevant number hoping to distract from the issue they can’t or don’t want to solve. If eating glue is just one of seven million unique queries but it’s searches by a million people, that’s more important than one of 7 thousand only searched by one person. The measure should be impact of bad queries and bad info. Or amount of garbage responded. Or even better, confusion created in people based on AI responses. They might be able to proxy this because they have chrome usage and android usage data and see what people do after these queries. Do they stop searching? Do they watch a movie? Do they jump off a building? Do they never search again (ie, die)?
- programjames 2y ago> AI will lead to more queries because I search differently with AI results than an index/page rank result. If you do 1000 queries a day, it would take you twenty years before you expect to hit a problematic query, assuming such errors are random. First, almost no one makes that many queries a day, and second, the errors come mostly from being silly, not when trying to find accurate information. > Also, it’s not about the unique queries, but the number of users running them. This seems like “data waving” where google is giving some not relevant number hoping to distract from the issue they can’t or don’t want to solve. I actually want to know once per 7M unique queries more than 10M users (or w/e the number is). Memes spread, so users don't tell me what their failure rate actually looks like. > The measure should be impact of bad queries and bad info. Or amount of garbage responded. Most stuff on the internet already is garbage advice and bad info. Much of this is due to SEO spam, but it's also because there are far fewer experts than bloggers (e.g. on the topic of AI). Now, I think the best solution would be for Google to start punishing SEO spam again so I can actually search for the right information by myself, but it seems possible that the AI Overview could be helpful in the sifting process. > Or even better, confusion created in people based on AI responses. I mean, this is a pretty difficult measure to figure out, even through polling. People aren't really trained to notice when they're confused, and it's pretty difficult to know you saw misinformation when you were looking up the information for the first time.
- jsemrau 2y ago“nobody asked the question” That's such an odd thing to say for a search engine operator.
- jsemrau 2y agoThey summarizer model has access to their core web page ranker. Since we saw Reddit and Onion posts in their result set for the summarizer, I'd argue that the goals of the core web page ranker don't align well with the goals of the summarizer.
- ADeerAppeared 2y ago> but Google seems to be pretty honest here. They're not. Take "This means that AI Overviews generally don't “hallucinate” or make things up in the ways that other LLM products might." This is just a lie, in two ways: 1. It's still an LLM, feeding search results into the prompts and asking for a summary reduces the probability of hallucination compared to directly asking the question in the prompt, it's still a probability non-negligibly above zero. This kind of pretending you can "fix" the hallucination problem by feeding data into the prompt is extremely wrong and dangerous. 2. It's missing the entire point. "Um akshually the LLM didn't hallucinate, we simply gave it Reddit user Fucksmith's post which it confidently restated as truth" (or indeed, the much funnier, "we gave it The Onion, which it confidently restated as truth") is functionally the same as a hallucination to the end user. The mechanism here does not matter. These are both fundamental critical issues to AI overviews. "Rare" hallucinations are unacceptable on a tool used billions of times a day. And the LLM paraphrasing lies/satire/ignorance/etc into truth is a fundamental flaw of using LLMs to paraphrase like this. You can't just patch this issue, the entire thing is fundamentally flawed.
- Ukv 2y ago> This is just a lie [...] it's still a probability non-negligibly above zero. This kind of pretending you can "fix" the hallucination problem [...] I don't think "generally don't hallucinate" and "it's usually for other reasons" are implying a total fix. From what I've seen, it is true that these errors stemmed from uncritically accepting satire/unreliable sources rather than hallucinations. > It's missing the entire point. "Um akshually the LLM didn't hallucinate, we simply gave it Reddit user Fucksmith's post which it confidently restated as truth" (or indeed, the much funnier, "we gave it The Onion, which it confidently restated as truth") is functionally the same as a hallucination to the end user. The mechanism here does not matter. When you can check the provided source it's summarising and see that it's a satire website, I think that's materially different than ChatGPT's hallucinations from thin air. It is still an issue, and if the post was just deflecting by saying "it's not typical hallucination so it doesn't matter" then I'd agree that'd be insufficient, but it seems totally fine to clarify that it's a different issue and go on to describe how they're working on fixing that issue. > "Rare" hallucinations are unacceptable on a tool used billions of times a day I don't think 100.0% accuracy is a reasonable bar for anything, including humans. My benchmark would be roughly "people come away with the correct information the same or more often as they did before, with features like highlighted snippets".