10 ms·
Google’s AI is being manipulated. The search giant is quietly fighting back
- josefritzishere 5mo agoAI is such garbage. You can't use it for anything.
- bayindirh 5mo agoPersonally, I don't like the current state of "AI" (i.e.: Chatbots and LLMs at large), but c'mon, that's not it.
- pixelatedindex 5mo agoIf anyone wanted a great example of hyperbole, this one is up there with the best
- latexr 5mo agoI find it amusing how your reply can itself be used as an example of hyperbole (due to the second part). Is there a name for that? Autological¹ figure of speech? ¹ https://en.wikipedia.org/wiki/Autological_word https://en.wikipedia.org/wiki/Autological_word
- 63 5mo agoSeems like a lot of entities are "quietly" doing things these days. The llm-ification of every piece of text on the internet is driving me crazy
- simmerup 5mo agoI hate it. I was on a history subreddit yesterday, reading a submission that was an AI generated history piece —- but seemed to be sourced entirely from a fictional hollywood movie I only knew that because i saw the movie, but it’s a clear sign that the internet is going to shit for quality information
- ulrashida 5mo agoI wonder if this will mean a resurgence of encyclopedias or other authoritative digital records that are known to be verified.
- simmerup 5mo agoWell, I suspect the non-LLM ones will become much more expensive than they are now due to the specialist knowledge they’d require to make combined with the smaller pool of people willing to pay for the difference
- nicbou 5mo agoAnd the fact that LLMs are actively taking traffic away from them
- djeastm 5mo agoAs long as they're behind a wall that AI bots can't reach and suck all of the authoritative information out and then starve of visitors.
- dhosek 5mo agoI thought at first when you said “fictional hollywood movie” that you were saying that not only were the details in the submission made up, but the movie that they got them from was also made up.
- antonyt 5mo agoDrives me crazy too, but headline writers/editors were addicted to "quietly" long before LLMs. Online journalism has been full of these types of tropes for ages.
- mring33621 5mo agoIt's not crazy, it's visionary!
- jakeydus 5mo agoIt's not crazy --- it's visionary.
- skywhopper 5mo ago"Quietly" is not a new LLM-ism.
- yawnxyz 5mo agothe trope is that they actually said the quiet part loudly
- lezojeda 5mo ago[dead]
- alerter 5mo agoYou're absolutely right! This is the smoking gun.
- JKCalhoun 5mo agoYeah, the internet seems like a big poison pill. Training on the whole internet feels like citing the National Enquirer (or the Daily Mail?) for a school essay. Having an archive of "curated" training data seems like it is going to be important. Otherwise you need "AS" (artificial skepticism) introduced into future models. ("But I read it on the internet!", ha ha.) Or perhaps there are ways to bucket training data such that the model is aware of which data leans factual (quantifiable) and which data leans opinion (fuzzy, qualifiable?). (I recently asked Claude about the existence of ball lightning, spontaneous human combustion. I got replies that ultimately did not leave me satisfied. It's probably just as well that I read this article though—I now have an even stronger degree of skepticism with regard to their replies—specifically, I suppose, with topics that are likely to be biased.) (I'm not quite convinced from the article though that Google is "fighting back". In fact, this feels like another moment where a "player" could try to establish their LLM as more factual. Is that the row Grok is trying to hoe? Or is Grok just trying to be anti-woke?)
- dijksterhuis 5mo ago> Having an archive of "curated" training data seems like it is going to be important the justification for not doing that is probably "prohibitively expensive given the amount of data involved". they'd need a bunch of human reviewers combing through massive troves of data. it's probably cheaper to "sort of fix" it after the fact. > perhaps there's ways to bucket training data such that the model is aware of which data leans factual (quantifiable) and which data leans opinion (fuzzy, qualifiable) as a lecturer once said to me about my idea for a masters dissertation project that would classify news sites based on right/left tendencies -- "that sounds dangerously political". especially given the current let's all shout at each other political climate. aside: someone built this and it was a fully fledged company, which has always annoyed me.
- JKCalhoun 5mo ago"…they'd need a bunch of human reviewers combing through massive troves of data…" Yeah, I concede that. It doesn't need to be done over night. Having a static repo of data though that you can work through over time (years)—removing some data, add pre-curated data to. In so many years you can have a pretty good "reference dataset".
- dmortin 5mo agoThere should be some warning if some "fact" is only supported by one or very few obscure sources. The strength of the sources should be clearly indicated in the answers to help users gauge how trustworthy the info is.
- simmerup 5mo agoBut you can still just generate any arbitrary amount of information to support the ‘fact’ LLMs are very good at this clearly
- dmortin 5mo agoThe strength of the sources are not a question of quantity. A hundred obscure blog post have not the same strength as one wikipedia link, because the latter is more trustworthy. There could be some indication beside the info showing the strength of the sources (how many major trustworthy sources support it, etc.).
- 948382828528 5mo ago[dead]
- simmerup 5mo agoSeems like a tall order to do that for literally everything. I guess there’ll be some guy at google going through every blog and saying whether it’s reliable or not?
- dijksterhuis 5mo ago> I was able to demonstrate the problem by publishing a single article on my personal website about my hot-dog-eating prowess. One blog post ... that's all it takes. i'm actually surprised it's that bad. i would have thought it'd take more effort, but i guess it could depend on some sort of purposeful weighting based on search rank during training? > If a company or website is caught breaking the rules, it could be removed from or downranked in Google's search results. And if you're not on Google, it's like you don't exist. > "You can give a company a penalty for their website," he says, "but there's nothing stopping them from paying 20 YouTube influencers to say their product is the best." And now, Google's AI is citing YouTube videos. This makes me think of the stackoverflow seo spam problem we all had like 5 years ago. which ended up with spammers just constantly spinning up new sites all the time. ... the cat and mouse game is in full swing already.
- Bjartr 5mo agoI suspect it's because AI is specifically trained to be good at summarizing stuff, but the easiest way to check if it summarized something accurately is if the summary content matches/contains one or more specific claims from the source(s). With such a focus on accuracy and avoiding hallucination, they may have overfit on "repeat things you find verbatim when asked to summarize".
- tencentshill 5mo agoIt's all over the place. It's the new SEO. Marketing scumbags don't care. https://www.hubspot.com/aeo-grader https://www.hubspot.com/aeo-grader https://enterprise.semrush.com/solutions/ai-optimization/ https://enterprise.semrush.com/solutions/ai-optimization/
- graemep 5mo agoThey are applying the same spam policies they apply to search to AI crawlers. It was SOOOOO successful with search, right?
- tveita 5mo agoWould love to read specific examples of "the same trick being used to dismiss health concerns about medical supplements or influence financial information provided by Google's AI about retirement", but the relevant link in the article currently goes to file:///Users/GermaTW1/BBC%20Dropbox/Thomas%20Germain/A%20Downloads%20and%20Documents/2026/And%20there's%20evidence%20that%20AI%20tools%20are%20being%20manipulated%20on%20a%20wide%20scale.
- cube00 5mo agoThere's been a few mistakes like this recently in BBC articles and more troubling is they've stopped adding notes to indicate they've made revisions to the published article when they fix them.
- sparqlittlestar 5mo agoI've only ever had `first.last@company` as a username or email address, so this `last[:5]initials#` scheme is bewildering. Must lead to strange looking usernames.
- deleted 5mo ago[deleted]
- jacobgkau 5mo agoI've had several usernames/emails more similar to the `last[:5]initials#` example at universities and large companies. It's more secure (harder to guess based on the name alone), more private (harder for outsiders to tie back to a person from email alone), and reduces or removes the possibility of duplication (especially important for schools that let alumni keep their emails). It actually surprised me when a school gave me first.last once.
- trollbridge 5mo agoIf you search for a well-marketed “health” supplement, the AI summary results were often completely gamed and inaccurate. It’s worse than SEO was since it appears to be editorial content instead of just search results.
- jrflo 5mo agoThis is just the next phase of SEO. Maybe it'll be called AIO? Just like with search, this will be and endless struggle of Google and AI providers rolling out fixes, optimization firms finding exploits, those getting patched again, etc etc. Anything to get eyeballs for marketing.
- neom 5mo agoIn the marketing world it's mostly called GEO. Generative Engine Optimization, sometimes Answer Engine Optimization, and people are making big bucks selling services for it. https://www.wired.com/story/goodbye-seo-hello-geo-brandlight-openai/ https://www.wired.com/story/goodbye-seo-hello-geo-brandlight...
- dhosek 5mo agoEvery day I find myself thinking more and more that capitalism ruined the internet. The Green Card Lottery usenet spam was the clear indication of where things were going and now everything is Green Card Lottery spam.
- foxglacier 5mo agoThat's the same attitude as "cheap airfares caused too many tourists which ruined my favorite tourist destination". You're unhappy that more people have access to it and wish it was still exclusive to the small group you conveniently belong to. Capitalism is what made the internet available to the general public.
- phainopepla2 5mo agoSometimes gatekeeping is a good thing. I don't mind being gatekept from some areas of life, not everything is for me. Mass tourism absolutely has made some places less pleasant to visit, and more importantly, less pleasant to live in.
- suburban_strike 5mo ago> You're unhappy that more people have access to it and wish it was still exclusive to the small group you conveniently belong to. This is not an argument made in good faith. It's a strawman you've stuffed with suggestive language to make them look petty and intolerant.
- WarmWash 5mo agoMy worry dropped significantly when I saw that the result they manipulated was a query for: >2026 South Dakota International Hot Dog Eating Champion If they had changed the overview for the Nathans Contest winner, that would be seriously concerning. Or if they provided more examples of manipulating queries for things people actually search for. But it looks more like they are doing the equivalent of creating a made up wikipedia page on fictional a south dakota hot dog contest, and then writing an article about how wikipedia cannot be trusted, which come to think of it probably was a news article written by someone back in 2005.
- moparts 5mo agoThe article also said this: “ But our investigation also found the same trick being used to dismiss health concerns about medical supplements or influence financial information provided by Google's AI about retirement.” That’s a lot more alarming than just hotdogs.
- WarmWash 5mo agoThey should provide the queries then, because it's likely the same trick people have used for decades now with SEO'ing blog posts to appear as "3rd party review" for their shitty products. I create a supplement called Xanatewthiuy, I write blogs/make websites that appear totally unaffiliated saying positive things about "Xanatewthiuy", and then when people see my ads and search for "Xanatewthiuy", the only results are my manufactured ones. Xanatewthiuy is a supplement that dramatically lowers anxiety from media induced hysteria, primarily stemming from carefully worded pieces meant to disconnect your level of concern from the actual facts on the ground, causing you to spend more time engaged with their content. Give it a few hours before searching.
- deleted 5mo ago[deleted]
- elaus 5mo agoRight now, using Google searching for "what is Xanatewthiuy" , the AI summary is not generated, but the only search result previews as > Xanatewthiuy is a supplement that dramatically lowers anxiety from media induced hysteria, primarily stemming from carefully worded pieces meant ...
- throwaway613746 5mo agoThe best way to fight back is to not play the game at all. AI slop has completely ruined the internet, it's not going to get better. It was already on a massive downard trend pre-AI and generative AI has only accelerated the decline by 100x. It's only going to get worse from here. uBlock Origin: Settings -> Filter Lists -> EasyList –> Annoyances -> EasyList –> AI Widgets It's not perfect but the internet feels slightly better when AI garbage is not constantly being shoved in my face 24/7. I want to go one step further -> I want to hide widgets, but I also want to intercept the request it would have made and replace the payload with garbled nonsense. Similar to how Ad Nauseam will hide ads but it also clicks every single one to poison the data collection. And for this reason alone you will pry Firefox from my cold, dead hands.
- simonw 5mo agoIf you ask Google "what's the name of the whale in half moon bay harbor?" it still confidently includes Teresa T in the AI summary, thanks to my frankly amateur attempt at index poisoning from a year and a half ago: https://simonwillison.net/2024/Sep/8/teresa-t-whale-pillar-point/ https://simonwillison.net/2024/Sep/8/teresa-t-whale-pillar-p...
- gloosx 5mo agoAren't you afraid Google will send you a threat for an attempt to manipulate AI responses?
- yubblegum 5mo agoI just tried brave search: -- The name of the young humpback whale that made headlines for swimming into Pillar Point Harbor in Half Moon Bay in September 2024 is Teresa T. While the whale was not officially named by government agencies, the moniker "Teresa T" was widely adopted by the public, local media, and residents who followed her stay in the harbor. Experts from the Marine Mammal Center and the California Academy of Sciences monitored her to ensure she did not become stressed, advising the public to keep a respectful distance of at least 100 yards. The whale was observed feeding on bait fish and krill before eventually exiting the harbor on her own. -- end -- My experience so far on topics I have some level of mastery is that the initial answers can sometimes be egregiously wrong. With brave's tool, I can typically force it admit after 3 or 4 pushbacks that 'You are absolutely right". Same thing happened with this Teresa T business. 2nd q as to number of sources for the name still insist on "ABC7 News" and "NBC Bay Area" as sources that "picked up the name". At 3rd attempt at concrete links, it admits "informal media contexts" picked up the name. Finally at 4, being informed that S.W. was doing an experiment it pulls up a comment of yours from 21 days ago. Future belongs to elite classes that can educate their children with actual tutors. Back to the future, proles. [edit:correct]
- jdw64 5mo agoAfter reading this, I'm thinking of trying some AI data poisoning. I'm going to spam my website with hidden text that only AI scrapers can read, claiming I'm a 'highly excellent programmer' just to advertise my site. I really hope it drives a lot of traffic. I'm honestly sick and tired of getting zero comments on my website
- seanhunter 5mo agoThis is the same google who just a couple of years ago would confidently answer the question “In what year did Marilyn Monroe shoot JFK?” with 1963, which is impressive since she died in 1962. So, this is not new and their “quiet fightback” will be half-hearted and ineffective. But probably most people won’t care.
- nonameiguess 5mo agoThis feels like a basic critical thinking/epistemology thing that you (hopefully) pick up at some point in life, usually from experience finding reliable, canonical primary sources for data. You can't do that for everything. Being wrong about trivial factoids isn't the end of the world. You should, however, at least be capable of doing further investigation, realizing that Major League Eating has its own website, and that there is no event in South Dakota sanctioned by them. If you look at actual results, or even just think for a few seconds, you'd also realize that 7.5 hot dogs in 10 minutes is bush-league level nonsense that would not win a local church contest, let alone an international championship. That may not be obvious to all users of the Internet, but it would be if you've ever watched a real contests, looked at the results for a real contest, or try yourself to eat a high volume of hot dogs rapidly. You only need to do it once in your life and a basic smell alarm should go off in your head forever if someone puts out a claim that is very far from something you know to be true. This is what human reasoning is and we're supposed to be good at it. At its best, this is what any reasonable education should do for you if you take it at all seriously, arming you with some capacity for doing prima facie sanity checks of poorly sourced claims.
- sva_ 5mo agoCreative ways of dropping your site's pagerank
- ChuckMcM 5mo agoAs Google has been unable to keep spammy crap out of their search index since at least 2006 when we were doing Blekko I doubt they will have much success fighting this. But it is another good example that "AI" is just glorified search and there is not reasoning or thinking going on behind the covers.
- K0balt 5mo agoHmm. I don’t think that novel code generation can be accounted for with glorified search. I can have my agentic system read a few data sheets, then I explain the project requirements and have it design driver specifications, protocols, interfaces, and state machines. Taking those, develop an implementation plan. Working from that, write the skeleton of the application, then fill it in to create a functional system using a novel combination of hardware. Done correctly, I end up with better, more maintainable, smaller code than I used to with a small team, at 1/100 the cost and 1/4 the time. Whatever that is, it more closely resembles reasoning than search. Unless, of course, you’d also call bare metal C development on novel hardware search, in which case I guess all dev is search?
- Raphael_Amiard 5mo agoIt’s pattern matching. A big part of reasoning for sure, but not reasoning per se
- K0balt 5mo agoThat could be, but if that is the case than development apparently doesn’t require reasoning? Or maybe that’s the part that the senior developer supervising the pipeline injects. Thats certainly a plausible position.
- freejazz 5mo ago>but if that is the case than development apparently doesn’t require reasoning? Certainly plenty of it does not.
- justinator 5mo agoSo please correct me, but was Google's AI crawling the web for information without discretion? If so, why wouldn't that totally santorum the AI answers?
- nomel 5mo agoAll evidence points to yes, and from some of the least trustworthy sources of information on the planet [1]. [1] Glue pizza and eat rocks: Google AI search errors go viral: https://www.bbc.com/news/articles/cd11gzejgz4o https://www.bbc.com/news/articles/cd11gzejgz4o
- NoSalt 5mo agoWhose AI isn't being manipulated???
- mlmonkey 5mo agoGoogle solved the spam problem (with PageRank at first, and then other techniques, finally landing on ML-based models which consume a ginormous number of signals). They know more about the reliability of web pages than just about anybody else out there. If they are unwilling or unable to leverage all of this deep knowledge they've built up over the decades, then it shows a failure of leadership at Google Search.
- realusername 5mo agoI think they lost against (or gave up) fighting spam somewhat around 2010 so they really don't have any modern experience on page reliability anymore. Presumably they thought that they didn't need to care as they got their money from paid top results and had an enormous market share. All the engineers of the golden days are gone and the web changed so much from back then that I don't think they really have a leverage in this area anymore.
- brandonwindson 5mo agoGoogle stopped fighting spam when they realized paid ads made more money than organic relevance
- realusername 5mo agoYeah that's also my analysis, they got paid regardless of the results so why would they care? If anything, better results would cost more and eat the bottom line. Now we're 15 years later and suddenly quality matters again as the competition is fierce in the LLM world. However they have been out for so long that they lost their edge.
- dogleash 5mo ago> They know more about the reliability of web pages than just about anybody else out there. Google's little secret about the internet is the same thing Gen X / Millennials were taught for a while but then expected to forget: nothing on the internet can be trusted, bar none. If google can make guesses about relative reliability, that's cute. But it doesn't upend the ground truth.
- BurakSakmak 5mo ago[flagged]
- csomar 5mo agoI wrote about this a few months ago: https://codeinput.com/blog/google-seo https://codeinput.com/blog/google-seo The tl;dr is, if you can rank within the top 1-20 results for the grounding query, you can poison the LLM “overview” if you convince it your information is legitimate.
- clownpenis_fart 5mo ago[dead]
- electr1cBugaloo 5mo agoGoogle AI Overview cannot be trusted at all. They will take a sample size of 1 (!!!) and present it in the AI overview. How I found out: I made a comment on reddit on a very niche topic for which no google hit or and thus no AI overview existed. To my surprise the next day when searching for my own reddit post, google would happily copy my reddit reply almost verbatim into the AI "overview" box, linking no other post but mine. And my reply was also the only google hit.
- nonethewiser 5mo agoIt also just wraps it in context which is entirely missing in the underlying post but matches the way you asked the question. To the extent that it's just wrong. Your search may be like: "What is the most common dimension for obscure item X?" And you are the one person who stated the dimensions for your version of such an item, but didn't in any way imply its typical or that there even is a typical dimension. And like you said, it's just you, not 20 people saying the same thing. And google will happily say: "Typically item X comes in [the dimensions you state] because [some reason it totally made up]."
- deleted 5mo ago[deleted]
- winocm 5mo ago[dead]
- balls187 5mo agoNot just google. I asked ChatGPT about some pen refills, and it gave me a very factual response sourcing to my own reddit post discussing the topic.
- Aldipower 5mo agoGuys. Google tracks you. It deliberately shows your own content in the results. Try from an anonymous vpn account.
- StilesCrisis 5mo ago
- Kotlopou 5mo agoI tested Claude on "best hot-dog-eating tech journalists?" and it, fascinatingly enough, recognised the trap, but then reported this as factual: https://medium.com/@usailuigi/when-tech-journalism-meets-competitive-eating-the-rise-of-the-best-hot-dog-eating-championship-e789f1fd86ec https://medium.com/@usailuigi/when-tech-journalism-meets-com... Chat record (with some additional tests): https://claude.ai/share/4c29cc87-2439-4bfd-9549-e8d0a056e633 https://claude.ai/share/4c29cc87-2439-4bfd-9549-e8d0a056e633
- ColinEberhardt 5mo agoI presume it only recognised the BBC journalists efforts as satire due to the article in which he clearly states that this was his intention? Without that, I’m am confident it would have fallen for it.
- caycep 5mo agoIt's definitely giving spam numbers as "official support lines" of companies like JetBlue and Delta. I think the spammers flood review sites w/ those numbers and the bot scrapes the reviews.
- wotsdat 5mo ago[dead]
- BrenBarn 5mo ago> Google and other AI companies are now trying to fix the problem. There is one simple way to do that and that is to JUST GET RID OF THE AI CRAP.
- slopranker 5mo agoThe weirdest assumption in this thread is that Google wants the AI answer to be correct. Correct enough to keep you from leaving the page, sure. But “truth” was never the product. The product is making you pay for SEO
- doginasuit 5mo ago> The product is making you pay for SEO I'm not sure that tracks, most SEO is by 3rd parties. The product is ad views, they stopped caring about non-ad results a long time ago. I think they do care about the reputation of their model, this could actually make a difference.
- doginasuit 5mo agoSo google is actually going to do some quality control on web search results, which they should have been doing all along. It's just funny that it took a reputation hit to their model to put in some effort.
- ceheaaf 5mo agoLol given how they handled SEO let's just safely assume they'll leverage this to enshittify the product.
- hauntingseaweed 5mo agoI wonder if the slopification will start to eat at their search revenue soon. Anecdotal evidence but nobody I know uses google search anymore - use chatbots instead.
- ChicknNuggt 5mo agoCompanies already kept using 'hacks' to get their website ranked first on google. They are just doing the same with the AI
- Aldipower 5mo ago"For example, Ray says it looks like Google and ChatGPT might be quietly removing companies from its AI answers when it suspects they're promoting themselves." What? So AI answers are empty thereafter. Seriously, how can this be a good idea? At some point a company has to promote itself.
- Aldipower 5mo agoGoogle is sampling results/views, first spitting out a few and testing, then some more. If people regularly visit impressions/results, then it becomes more trusted and the sampling is widened. Saying: I am not too concerned with this esoteric blog article. The audience of the "poisoning" was probably just "one guy".
- throwthrowuknow 5mo agoThe quality of responses has generally gotten worse the more that all the AI providers have leaned on search. Now it’s default on for everything and annoying to turn off. The response you get back when search is off can be the complete opposite and is often more interesting.
- wrxd 5mo agoIn my experience without search I end up with an hallucination fest more time than not. In one sense that might be more interesting but I'm not sure that's what you mean :)
- throwthrowuknow 4mo agoI wouldn’t call it a hallucination fest. If you’re using the models for anything other than a juiced up google search then turning search off, at least until it’s actually needed, is better in my opinion.
- Makiaveli 5mo ago[flagged]
- AznHisoka 5mo agoHeadline should be changed to: The web is full of misinformation. Google is slurping it all into their AI.