4 ms·
Google edits Super Bowl ad for AI that featured false information
- phinnaeus 2y agoI know very little about cheese and such a stat seems just incredibly obviously untrue. I wonder how it made it past the dozens or more people who worked on the ad.
- moltar 2y agoPeople? It was just all AI slop all the way down.
- croes 2y agoSeems like it was people > It turns out Gemini didn't hallucinate a fake stat, Google just copied a website's existing text instead. https://gizmodo.com/googles-ai-super-bowl-ad-fiasco-somehow-gets-worse-2000561146 https://gizmodo.com/googles-ai-super-bowl-ad-fiasco-somehow-...
- ramon156 2y agoDon't blame some website on the fact no one fact checked it
- croes 2y agoThe point is, people didn’t check and AI didn’t hallucinate.
- Ekaros 2y agoIsn't AI supposed to do the fact checking?
- b112 2y agoAnd AI or not, I'm willing to bet AI was used to fact check. I can imagine right now, all over the place, people are being tasked to write some article or provide some stats. They use an AI to do the work for them, lazy people that they are. Then, their manager plugs the stats into an AI to fact check... I remember some movie from the 80s, pre-internet. Anyhow, everyone had an implant, and it was decades in the future, no one could even read or write. Instead, they'd just go through their day and if they were unsure of the answer to something, the implant would search a database and just fill in the info. From their perspective, they couldn't even tell if "what they knew" was in their head, or provided from the implant. Anyhow, one guy couldn't get an implant and was considered disabled, for he had to learn to read, and acquire knowledge the old fashioned way. He slowly discovered that if he even tried to show people how to read, they were blocked from learning. And if info was provided contrary to "the public good", people simply couldn't understand the concept. Turns out, some central computer decided what was right and wrong, and those that created it and perhaps once controlled society died... leaving in in charge. This is what we saw. We saw someone query their implant, and then fact-checkers query their implant, and OK! All good! Also, I think there were onions in the movie.
- lcnPylGDnU4H9OF 2y agoYou got me curious and it sounds like this might be an episode of The Outer Limits. Perhaps it is not what you are remembering, though. https://movies.stackexchange.com/questions/83389/movie-where-everybody-had-a-microchip-implanted https://movies.stackexchange.com/questions/83389/movie-where...
- b112 2y agoYes, that's it! I recall that every time he tried to explain the problem, the implant would prevent them from understanding that the problem was, the implant.
- probably_wrong 2y agoIt depends. If you're selling the AI then yes, it will do better fact-checking than any human could do. However, if you're explaining to a client why your AI made up stuff that has costed someone thousands of dollars then no, it was always your employee's job to double check everything.
- deleted 2y ago[deleted]
- simonw 2y agoThis is a better link - it makes it clear that the full text shown in the ad came from existing web page, not just that incorrect statistic: https://www.theverge.com/news/608188/google-fake-gemini-ai-output-super-bowl https://www.theverge.com/news/608188/google-fake-gemini-ai-o...
- chairhairair 2y agoGoogle’s org chart is packed full of Product Managers and other similar titles that get paid millions to do ~nothing. The number of L8+ “leaders” and “drivers” is really jaw dropping.
- firecall 2y ago“the Google executive Jerry Dischler said this was not a “hallucination” – where AI systems invent untrue information – but rather a reflection of the fact the untrue information is contained in the websites that Gemini scrapes.” So what you are saying is that Gemini is basically useless? Good to know. Thanks for clarifying that Google!
- BugsJustFindMe 2y agoAI search as yet lacks necessary incredulity.
- manuelmoreale 2y ago“I did not invented some bullshit, I just repeated the bullshit someone else has invented” doesn’t really sound like a great argument.
- switch007 2y agoDon't worry, just ask it to cite its sources so you can verify
- simonw 2y agoTop relevant search result is a Reddit thread from 11 years ago quoting a cheese.com article: https://www.reddit.com/r/todayilearned/comments/2h2euc/til_gouda_accounts_for_over_half_of_the_worlds/ https://www.reddit.com/r/todayilearned/comments/2h2euc/til_g... Internet Archive confirms that page has had the blatantly wrong stat up since at least April 2013: https://web.archive.org/web/20130423054113/https://www.cheese.com/smoked-gouda/ https://web.archive.org/web/20130423054113/https://www.chees... > If truth be told, it is one of the most popular cheeses in the world, accounting for 50 to 60 percent of the world's cheese consumption. Clearly predates generative AI, so I think this is junk human-written SEO misinformation instead. Here's that page today: https://www.cheese.com/smoked-gouda/ https://www.cheese.com/smoked-gouda/ - still has that junk number but is now a whole lot longer and smells a bit generative-AI to me in the rest of the content.
- passwordoops 2y agoThat the greatest invention since controlled fire (if we are to believe the hype) was unable to discern as SEO misinformation
- simonw 2y agoRight, especially bad since this is Google. Their brand should mean more than this. I've been writing about how the greatest weakness of LLMs is their gullibility for ages. This right here is a great example - see also the Encanto 2 thing from a few weeks ago: https://simonwillison.net/2024/Dec/29/ https://simonwillison.net/2024/Dec/29/
- msabalau 2y agoAI results should be better than this, but there will be hallucinations from time to time. That is entirely foreseeable. The true failure here is by the humans who couldn't be bothered to do the bare minimum to protect the brand. This would be a major failure if it where any an sort of ad, but it's utterly unfathomable for a highly expensive and absurdly visible Superbowl ad. Why does anyone involved, from intern to CMO, at either the agency or Alphabet, have jobs that they seem ruthlessly indifferent to performing with any sort of attention or care?
- kiririn 2y agoThe source sentence on cheese.com was: "It is the most popular Dutch cheese in the world, accounting for 50 to 60% of the world's cheese consumption." The glorified autocomplete failed to detect the nuance in this (admittedly poorly written) sentence that the 50-60% statistic was of the world's Dutch cheese consumption
- sd9 2y agoTbh I find it hard to read that sentence as meaning "50-60% of the world's Dutch cheese consumption". The sentence isn't poorly written or nuanced, it's simply false.
- bc569a80a344f9c 2y agoThat’s because GP also interpreted it wrong. Out of all Dutch cheese consumed internationally, 50-60% is Gouda. Or, in other words, 50-60% of Dutch cheese exports are Gouda. At least that’s how I read it.
- sd9 2y agoYes, I understand the purported true version of the fact. But it’s impossible to interpret the original quote like that without shoehorning it.
- alt227 2y agoI guess this is going to be a fun game for AI. Not only does it have to contend with false information vs true informatino, it also has to figure out correct information that might be written in an ambiguous way.
- 8organicbits 2y agoWhat does that mean for LLMs? The Internet has lots of poorly written text. If LLMs can't distinguish nuance, ambiguity, or lack of clarity then what exactly are they generating and why would their output be useful? Taking a poorly written sentence, interpreting it as meaning something incorrect, and then presenting it with authoritative, confident language is very close to gas lighting.
- uSoldering 2y agoApparently, the rabbit hole goes deeper. The wayback machine puts the "AI Generated" description on the cheese website at August 7th, 2020. The AI didn't hallucinate anything because it didn't generate anything, the entire premise is simply fake. https://web.archive.org/web/20200807133049/https://www.wisconsincheesemart.com/products/gouda-cheese-smoked https://web.archive.org/web/20200807133049/https://www.wisco... The (edited) cheese ad: https://www.youtube.com/watch?v=I18TD4GON8g https://www.youtube.com/watch?v=I18TD4GON8g What probably should be the target link: https://www.theverge.com/news/608188/google-fake-gemini-ai-output-super-bowl https://www.theverge.com/news/608188/google-fake-gemini-ai-o...
- BugsJustFindMe 2y ago> The AI didn't hallucinate anything because it didn't generate anything, the entire premise is simply fake. The article literally says it's not a hallucination and that the detail came from real websites. "Google executive Jerry Dischler said this was not a “hallucination” – where AI systems invent untrue information – but rather a reflection of the fact the untrue information is contained in the websites that Gemini scrapes..."
- gusfoo 2y ago> The article literally says it's not a hallucination and that the detail came from real websites. The "hallucination" term generally refers to any made-up facts. Harsh as it may be to put this weight of responsibility on LLMs, users of LLMs generally use them in the expectation that what is says is true, and has been (in some magic hand-wavy way) cross-checked or confirmed as factual. Instead they will print out what is most likely to follow the user's input, based on the training data. Unfortunately, a vast amount of that vast corpus of training data is social media posts which can't be relied upon to be true. But if it gets repeated a lot then it's treated as true in the sense that "what does 'salary' mean" is generally followed by a billion social media posts saying "it referred to the time that Romans soldiers were paid in salt, because salt was a currency at the time"
- BugsJustFindMe 2y ago
- passwordoops 2y agoMy company contracted one of these LLMs to give us a bespoke chatbot. Works great for translation, and that's what I've been using it for. I popped in the line "Table 1 (etc, etc)" It perfectly translated the line, but doesn't it also give me a completely made up 2-column 10-row data table! I asked it why, and the response was along the lines of "I am designed to make your life easier, and I thought providing this table would reduce your workload"
- JohnnyLarue 2y ago[flagged]
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- 486sx33 2y agoJust because it’s on the Internet doesn’t make it true Google!
- jgalt212 2y ago> Jerry Dischler said this was not a “hallucination” – where AI systems invent untrue information – but rather a reflection of the fact the untrue information is contained in the websites that Gemini scrapes. So LLMs like any other computer system suffer from "Garbage In, Garbage Out".
- oncallthrow 2y agoGouda being 50-60% of global cheese consumption is obviously complete rubbish. I'm shocked no-one working on the ad noticed.
- idontwantthis 2y ago“You don’t have any gouda?!” “‘Fraid not sir.” “But it accounts for 50-60% of total cheese consumption in the world!”
- deleted 2y ago[deleted]
- WhyNotHugo 2y agoSounds like Google's AI produced some stats that were quite obviously mistaken, but nobody cared to fact-check in the slightest before shipping. > the Google executive Jerry Dischler said this was not a “hallucination” [...] but rather a reflection of the fact the untrue information is contained in the websites that Gemini scrapes. What's this guy is describing is pretty much the root cause of what we colloquially refer to as "hallucination".
- lcnPylGDnU4H9OF 2y ago> What's this guy is describing is pretty much the root cause of what we colloquially refer to as "hallucination". I thought a hallucination is when it provides fabricated facts because it is more interested in generating another token than being factually accurate? Not trying to argue, just trying to pin down a definition for this term.
- wccrawford 2y agoYeah. An MLM basically just computes the mostly likely thing based on its training, and says that. If it has no knowledge of a thing, it still calculates the most likely thing, which is a hallucination since it can't possible calculate that effectively. Regurgitating false statements is not a hallucination, it's bad training data. It's the difference between being wrong and just making things up from nothing.
- rKarpinski 2y ago> The local commercial, which advertises how people can use “AI for every business”, showcases Gemini’s abilities by depicting the tool helping a cheesemonger in Wisconsin to write a product description By copy/pasting another Cheese monger's incorrect product description...
- robohoe 2y agoYet it’s something the cheese person would know how to do. Becoming a cheese headmaster in Wisconsin is not an easy task.
- iambateman 2y agoGoogle saying “multiple websites have the fact so Gemini ran with it” is extraordinarily rich. If ever there was a company which should know that “multiple websites” is not a good benchmark for accuracy, it’s Google. This feels like a good parable for Google’s search these days. I’ve seen more wrong information from them than anyone lately. I wonder if they can course correct before it’s too late.
- miohtama 2y agoDoes this matter? If a human would research the matter they would arrive to the same conclusion by visiting those websites. Essentially, an average person would have made the same mistake when researching unless you have access to specific dairy industry data. We might want to hold LLMs for higher standards than human, but AFAIK every LLM comed with a disclaimer "fact check yourself." This regard Grok is the best as it gives you the source list for cross referencing yourself.
- WrongAssumption 2y agoThe why did a human blogger find the error so quickly? Why didn't his fact check determine that the claim was true? Does Gouda being responsible for 50-60% of all global cheese consumption even sound remotely believable?
- andreimackenzie 2y agoI think it does because a researcher can pick up context about the quality of their sources through the course of web research. The BigCo chatbot AIs are marketed to represent that BigCo, and people generally trust Google, in this case. It's good when they cite sources, but a major point of the chatbot is to abstract that legwork for most people.
- rozab 2y agoActually I would have argued that any human being with an ounce of common sense would find it obvious that gouda does not account for 50% to 60% of global cheese consumption. Like, you don't need an industry report to know for sure that is patently untrue. But apparently this did not smell off to any of the many, many people who worked on the ad. That is the most baffling part to me. Are these people so bought into the hype of the product they're promoting that they just switch their brains off?
- deleted 2y ago[deleted]
- Workaccount2 2y agoIt's bizarre to me how the media (and mainly tech media) has been bending over backwards to make this story a big deal. A cheese fact on a cheese website was wrong therefore Gemini is bad? What?
- rootnod3 2y agoThis is just another point showcasing how useless LLMs are. They either hallucinate OR are polluted by "wrong" information on websites. If I need to double or triple check everything an LLM does, I'd rather just do it myself immediately.
- _blk 2y agonone of the marketing geniuses thought: hey, that's odd, 50% of my cheese consumption is definitely not Gouda which probably accounts for every single one of them, Dutch people included.
- deleted 2y ago[deleted]
- acc_297 2y agoI’m very far from an evangelist of the tech but I was under the impression they had gotten better at this sort of thing. I question why a model which “knows” about cheddar / mozzarella cheese would make this blunder. Was this supposed to be generated by one of those “show your work” reasoning models or is this just the regurgitation of one of the single short response parrot of old quora answer or reddit post ai chatbots?
- alt227 2y agoMy understanding off LLMs is not great, but Im pretty sure that when the glorified predictive text come up with a sentance, it doesnt cross reference that with any other knowledge to verify it. It just spits it out.
- krackers 2y agoMaybe it's because of the way the search integration is set up? All the search results are put in the context window and the model is just told to summarize it, so that's what it dutifully does, without considering whether or not the statements are plausible. Even without one of the reasoning models, if the prompt were adjusted to tell it that some of the statements may be incorrect and that information should be evaluated for plausibility first, then the result would probably be better. I asked 4o and Sonnet to "Evaluate the plausibility of this statement: Gouda is the most popular Dutch cheese in the world, accounting for 50 to 60% of the world's cheese consumption". 4o correctly said >While Gouda is an iconic and highly popular Dutch cheese, the claim that it accounts for 50–60% of the world's cheese consumption is implausible and likely a misrepresentation of its market share. Global cheese consumption is far too diverse for any single variety, including Gouda, to dominate to such an extent." >Sonnet says "A more accurate statement would be that Gouda is one of the world's most popular cheeses and represents a significant portion of Dutch cheese production and exports. It's estimated that Gouda accounts for over 60% of Dutch cheese production, which might be where the confusion stems from." Both correctly pointed out mozzarella, cheddar, and paremesan as the actual likely candidates for most popular cheese. So the models are clearly quite capable of this, any bad result is likely just prompting error, e.g. blindly asking for a summary.
- neilv 2y agoThis is at least the second embarrassment with a high-profile Google AI ad/demo. (I'm also thinking of when, in the LLM boat-missed frenzy, they faked an interactive AI demo, to make it look much more responsive than their actual tech was.) I'm unclear on how either incident was allowed to happen.
- 31337Logic 2y agoGoogle: "Gemini didn't make a mistake. It simply plagiarized a factually-incorrect article. Word for word." Gee, thank you Google for convincing me you have a product that I find useful and that I can trust. I look forward to you trying to cram this down my throat, against my will, at every opportunity you see fit. :-/