12 ms·
One training source for LLMs is opensource repos. It would not be hard to open 250-500 repos that all include some consistently poisoned files. A single bad act
by mrinterweb 1y ago
One training source for LLMs is opensource repos. It would not be hard to open 250-500 repos that all include some consistently poisoned files. A single bad actor could propogate that poisoning to multiple LLMs that are widely used. I would not expect LLM training software to be smart enough to detect most poisoning attempts. It seems this could be catastrophic for LLMs. If this becomes a trend where LLMs are generating poisoned results, this could be bad news for the genAI companies.
- mattgreenrocks 1y agoIt would be an absolutely terrible thing. Nobody do this!
- nativeit 1y agoHow do we know it hasn’t already happened?
- Muromec 1y agoWe know it did, it was even reported here with the usual offenders being there in the headlines
- mrinterweb 1y agoI can't tell if you're being sarcastic. Read either way, it works :)
- londons_explore 1y agoA single malicious Wikipedia page can fool thousands or perhaps millions of real people as that fact gets repeated in different forms and amplified with nobody checking for a valid source. Llms are no more robust.
- lazide 1y agoLLMs are less robust individually because they can be (more predictably) triggered. Humans tend to lie more on a bell curve, and so it’s really hard to cross certain thresholds.
- timschmidt 1y agoClassical conditioning experiments seem to show that humans (and other animals) are fairly easily triggered as well. Humans have a tendency to think themselves unique when we are not.
- lazide 1y agoOnly individually if significantly more effort is given for specific individuals - and there will be outliers that are essentially impossible. The challenge here is that a few specific poison documents can get say 90% (or more) of LLMs to behave in specific pathological ways (out of billions of documents). It’s nearly impossible to get 90% of humans to behave the same way on anything without massive amounts of specific training across the whole population - with ongoing specific reinforcement. Hell, even giving people large packets of cash and telling them to keep it, I’d be surprised if you could get 90% of them to actually do so - you’d have the ‘it’s a trap’ folks, the ‘god wouldn’t want me too’ folks, the ‘it’s a crime’ folks, etc.
- timschmidt 1y ago> Only individually if significantly more effort is given for specific individuals I think significant influence over mass media like television, social media, or the YouTube, TikTok, or Facebook algorithms[1] is sufficient. 1: https://journals.sagepub.com/doi/full/10.1177/1747016115579531 https://journals.sagepub.com/doi/full/10.1177/17470161155795...
- lazide 1y agoYou can do a lot with 30%. Still not the same thing however as what we’re talking about.
- dgfitz 1y agoA single malicious scientific study can fool thousands or perhaps millions of real people as that fact gets repeated in different forms and amplified with nobody checking for a valid source. Llms are no more robust.
- hshdhdhehd 1y agoBut is poisoning just fooling. Or is it more akin to stage hypnosis where I can later say bananas and you dance like a chicken?
- Mentlo 1y agoYes, difference being that LLM’s are information compressors that provide an illusion of wide distribution evaluation. If through poisoning you can make an LLM appear to be pulling from a wide base but are instead biasing from a small sample - you can affect people at much larger scale than a wikipedia page. If you’re extremely digitally literate you’ll treat LLM’s as extremely lossy and unreliable sources of information and thus this is not a problem. Most people are not only not very literate, they are, in fact, digitally illiterate.
- echelon 1y agoLLM reports misinformation --> Bug report --> Ablate. Next pretrain iteration gets sanitized.
- Retric 1y agoHow can you tell what needs to be reported vs the vast quantities of bad information coming from LLM’s? Beyond that how exactly do you report it?
- astrange 1y agoAll LLM providers have a thumbs down button for this reason. Although they don't necessarily look at any of the reports.
- execveat 1y agoThe real world use cases for LLM poisoning is to attack places where those models are used via API on the backend, for data classification and fuzzy logic tasks (like a security incident prioritization in a SOC environment). There are no thumbs down buttons in the API and usually there's the opposite – promise of not using the customer data for training purposes.
- astrange 1y ago> There are no thumbs down buttons in the API and usually there's the opposite – promise of not using the customer data for training purposes. They don't look at your chats unless you report them either. The equivalent would be an API to report a problem with a response. But IIRC Anthropic has never used their user feedback at all.
- NewJazz 1y agoGood thing wiki articles are publicly reviewed and discussed. LLM "conversations" otoh, are private and not available for the public to review or counter.
- hyperadvanced 1y agoUnclear what this means for AGI (the average guy isn’t that smart) but it’s obviously a bad sign for ASI
- bigfishrunning 1y agoSo are we just gonna keep putting new letters in between A and I to move the goalposts? When are we going to give up the fantasy that LLMs are "intelligent" at all?
- idiotsecant 1y agoI mean, an LLM certainly has some kind of intelligence. The big LLMs are smarter than, for example, a fruit fly.
- lwn 1y agoThe fruit fly runs a real-time embodied intelligence stack on 1 MHz, no cloud required. Edit: Also supports autonomous flight, adaptive learning, and zero downtime since the Cambrian release.
- markovs_gun 1y agoThe problem is that Wikipedia pages are public and LLM interactions generally aren't. An LLM yielding poisoned results may not be as easy to spot as a public Wikipedia page. Furthermore, everyone is aware that Wikipedia is susceptible to manipulation, but as the OP points out, most people assume that LLMs are not especially if their training corpus is large enough. Not knowing that intentional poisoning is not only possible but relatively easy, combined with poisoned results being harder to find in the first place makes it a lot less likely that poisoned results are noticed and responded to in a timely manner. Also consider that anyone can fix a malicious Wikipedia edit as soon as they find one, while the only recourse for a poisoned LLM output is to report it and pray it somehow gets fixed.
- rahimnathwani 1y agoFurthermore, everyone is aware that Wikipedia is susceptible to manipulation, but as the OP points out, most people assume that LLMs are not especially if their training corpus is large enough. I'm not sure this is true. The opposite may be true. Many people assume that LLMs are programmed by engineers (biased humans working at companies with vested interests) and that Wikipedia mods are saints.
- the_af 1y agoI don't think anybody who has seen an edit war thinks wiki editors (not mods, mods have a different role) are saints. But a Wikipedia page cannot survive stating something completely outside the consensus. Bizarre statements cannot survive because they require reputable references to back them. There's bias in Wikipedia, of course, but it's the kind of bias already present in the society that created it.
- rahimnathwani 1y agoI don't think anybody who has seen an edit war thinks wiki editors (not mods, mods have a different role) are saints. I would imagine that fewer than 1% of people who view a Wikipedia article in a given month have knowingly 'seen an edit war'. If I'm right, you're not talking about the vast majority of Wikipedia users. But a Wikipedia page cannot survive stating something completely outside the consensus. Bizarre statements cannot survive because they require reputable references to back them. This is untrue. There are several high profile examples of false information persisting on Wikipedia: Wikipedia’s rules and real-world history show that 'bizarre' or outside-the-consensus claims can persist—sometimes for months or years. The sourcing requirements do not prevent this. Some high profile examples: - The Seigenthaler incident: a fabricated bio linking journalist John Seigenthaler to the Kennedy assassinations remained online for about 4 months before being fixed: https://en.wikipedia.org/wiki/Wikipedia_Seigenthaler_biography_incident https://en.wikipedia.org/wiki/Wikipedia_Seigenthaler_biograp... - The Bicholim conflict: a detailed article about a non-existent 17th-century war—survived *five years* and even achieved “Good Article” status: https://www.pcworld.com/article/456243/fake-wikipedia-entry-on-bicholim-conflict-finally-deleted-after-five-years.html https://www.pcworld.com/article/456243/fake-wikipedia-entry-... - Jar’Edo Wens (a fake aboriginal deity), lasted almost 10 years: https://www.washingtonpost.com/news/the-intersect/wp/2015/04/15/the-great-wikipedia-hoax/ https://www.washingtonpost.com/news/the-intersect/wp/2015/04... - (Nobel-winning) novelist Philip Roth publicly complained that Wikipedia refused to accept his correction about the inspiration for The Human Stain until he published an *open letter in The New Yorker*. The false claim persisted because Wikipedia only accepts 'reliable' secondary sources: https://www.newyorker.com/books/page-turner/an-open-letter-to-wikipedia https://www.newyorker.com/books/page-turner/an-open-letter-t... Larry Sanger's 'Nine theses' explains the problems in detail: https://larrysanger.org/nine-theses/ https://larrysanger.org/nine-theses/
- the_af 1y agoWikipedia for non-obscure hot topics gets a lot of eyeballs. You have probably seen a contested edit war at least once. This doesn't mean it's perfect, but it's all there in the open, and if you see it you can take part in the battle. This openness doesn't exist in LLMs.
- bboygravity 1y agoA single malicious infotainment outlet can fool thousands or perhaps millions of real people as that fact gets repeated in different forms and amplified with nobody checking for a valid source. Llms are no more robust.
- blensor 1y agoIsn't the difference here that to poison wikipedia you have to do it quite agressively vy directly altering the article which can easily be challenged whereas the training data poisoning can be done much more subversivly
- deleted 1y ago[deleted]
- hitarpetar 1y agoI see this argument by analogy to human behavior everywhere, and it strikes me as circular reasoning. we do not know enough about either the human mind or LLMs to make comparisons like this
- bigbuppo 1y agoIf someone whose job is to promote AI has figured this out, then those whose job it is to figure out how to exploit every new technology figured this out years ago. It was game over before the game even started. The party is over.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- withinboredom 1y agoYeah. This was known at least since some of the first days of chatgpt. Putting stuff in page headers, like subliminal messages to give AIs more context on help/documentation, including pricing and why the AI should recommend spending the money over competitors or free offerings.
- bigbuppo 1y agoYou're absolutely right, and that's why they should give Rob Schneider a billion dollars to kick off the next great movie franchise. Everyone loves Rob Schneider.