4 ms·
I think the difference between our takes is our understanding of how the "AI Overview" was trained. Google said, > AI Overviews work very differently than chat
by programjames 2y ago
I think the difference between our takes is our understanding of how the "AI Overview" was trained. Google said,
> AI Overviews work very differently than chatbots and other LLM products that people may have tried out
and later,
> Let’s take a look at an example: “How many rocks should I eat?” Prior to these screenshots going viral, practically no one asked Google that question.
It seems it's being trained based on search queries. This is important to protect against confabulation, but obviously AI will fail on out-of-distribution data. Such as search queries that practically no one asked about before. Now, they could ask an LLM if the AI Overview seems reasonable, but the entire point of their alternate training scheme is to guarantee the output can be backed with sources. Why introduce a potential failure mechanism?
You say, "they gaslight us with words like 'faithful' and 'nobody asked the question,'" but given this is how their model works, it seems like the best, non-gaslighty defense. Plus, it's pretty unreasonable to expect Google to fix every potential bug before production. As they said,
> there’s nothing quite like having millions of people using the feature with many novel searches. We’ve also seen nonsensical new searches, seemingly aimed at producing erroneous results.
Having "a content policy violation on less than one in every 7 million unique queries" seems pretty solid to me. I'm sure it'd be much higher if they were using an LLM like Claude.
EDIT: My argument is pretty much, it looks like they designed this to be helpful and accurate, so the errors seem like honest mistakes from imperfect execution rather than deception or politicking. Their post seems pretty open about acknowledging what went wrong and that they want to do better.
- gmerc 2y agoSo they failed to make it work and are now trying to shift the goalposts. They should not have launched it then.
- prepend 2y ago> Having "a content policy violation on less than one in every 7 million unique queries" seems pretty solid to me. Not to me. AI will lead to more queries because I search differently with AI results than an index/page rank result. Also, it’s not about the unique queries, but the number of users running them. This seems like “data waving” where google is giving some not relevant number hoping to distract from the issue they can’t or don’t want to solve. If eating glue is just one of seven million unique queries but it’s searches by a million people, that’s more important than one of 7 thousand only searched by one person. The measure should be impact of bad queries and bad info. Or amount of garbage responded. Or even better, confusion created in people based on AI responses. They might be able to proxy this because they have chrome usage and android usage data and see what people do after these queries. Do they stop searching? Do they watch a movie? Do they jump off a building? Do they never search again (ie, die)?
- programjames 2y ago> AI will lead to more queries because I search differently with AI results than an index/page rank result. If you do 1000 queries a day, it would take you twenty years before you expect to hit a problematic query, assuming such errors are random. First, almost no one makes that many queries a day, and second, the errors come mostly from being silly, not when trying to find accurate information. > Also, it’s not about the unique queries, but the number of users running them. This seems like “data waving” where google is giving some not relevant number hoping to distract from the issue they can’t or don’t want to solve. I actually want to know once per 7M unique queries more than 10M users (or w/e the number is). Memes spread, so users don't tell me what their failure rate actually looks like. > The measure should be impact of bad queries and bad info. Or amount of garbage responded. Most stuff on the internet already is garbage advice and bad info. Much of this is due to SEO spam, but it's also because there are far fewer experts than bloggers (e.g. on the topic of AI). Now, I think the best solution would be for Google to start punishing SEO spam again so I can actually search for the right information by myself, but it seems possible that the AI Overview could be helpful in the sifting process. > Or even better, confusion created in people based on AI responses. I mean, this is a pretty difficult measure to figure out, even through polling. People aren't really trained to notice when they're confused, and it's pretty difficult to know you saw misinformation when you were looking up the information for the first time.