7 ms·
Can you simply brainwash an LLM?
- danbrooks 3y agoIs this surprising? LLMs are trained to produce likely word/tokens in a dataset. If you include poisoned phrases in training sets, you’ll surely get poisoned results.
- munchler 3y agoThey’re “surgically” corrupting an existing LLM, not training a new LLM with false information. This requires somehow finding and editing specific facts within the model.
- taneq 3y agoAh, so LoRA / RLHF?
- gmerc 3y agoThere’s a word for that: Finetuning. It’s a feature not an attack.
- ravi-delia 3y agoFine-tuning is usually used to specialize a model. In this case they were really trying to change a small aspect of behavior without altering performance on other tasks. It's not surprising that it worked or anything, but I'm not aware of anyone publishing something like this prior. They describe it as an attack because just looking at the weights there really isn't a way to tell if a model has had this sort of thing done to it- you're unlikely to notice the tweaked fact because on any other task it behaves identically. So someone could sneak things in with downstream users being none the wiser. What could you do with that? I can't think of anything. But it's apparently possible!
- deleted 3y ago[deleted]
- ryaclifton 3y agoDoes this mean that I could train an LLM to do something like spread fake news? Would that even scale?
- laverya 3y agoIsn't this done with every "sanitized" LLM? Fake news is all according to perspective!
- jdiff 3y agoNo, it isn't. This is akin to saying that the truth is relative and lies somewhere between "the Earth is an oblate spheroid" and "the Earth is flat." Perception and perspective varies, sure, but fact exists regardless. Fake news is falsified news built on fabricated fakes, and is not just alternative viewpoints. Do not normalize this.
- laverya 3y ago"Japan has a higher GDP per capita than Alabama" is fake news. It's also confidently repeated by most LLMs. https://twitter.com/MatthewJBar/status/1681554646664634368 https://twitter.com/MatthewJBar/status/1681554646664634368 My assertion is that things like this will happen whenever LLMs are tuned to match political beliefs.
- jdiff 3y agoNobody fed that false fact into any LLM. Garbage in, garbage out isn't fake news.
- sixothree 3y agoProbably. And you could surround specific communities en masse. And it’s coming soon to every single site near you.
- jacquesm 3y ago
- The28thDuck 3y agoI feel intuitively this makes sense. You can tell kids that cows in the South moo in a southern accent and they will merrily go on their way believing it without having to restructure their entire world view. It goes with the problem of “understanding” vs parroting. Human-centric example but you get the point.
- dTal 3y agoKids, but not adults. What's the difference? A more interconnected world model with underlying structure. LLMs have such structure as well, proportional to how well they're trained. A "stupid" model will be more easily convinced of a counterfactual than a "smart" one. And similarly, the limits of counterfactuality a child is prepared to believe is (inversely) proportional to their age.
- chrisnight 3y agoThere is a certain balance in this act though. Malleability of facts or opinions can be a sign of maturity and not youth. While the types of malleability for adults and young kids are different, with adults generally requesting evidence and reasonings before changing their mind, in the instance of an LLM, where it has no access to “evidence” other than what you tell it, it has to at some point accept what the user tells it if it wants to be the best it can. Otherwise you’ll get Bing Chat again with the “I don’t believe you.” responses to pure facts.
- davidguetta 3y agolol how many adults believe the earth is flat >< ?
- iambateman 3y agoThe people pushing this line of concern are also developing AICert to fix it. While I’m sure they’re right - factually tampering with an LLM is possible - I doubt that this will be a widespread issue. Using an LLM knowingly to generate false news seems like it will have similar reach to existing conspiracy theory sites. It doesn’t seem likely to me that simply having an LLM will make theorists more mainstream. And intentional use wouldn’t benefit from any amount of certification. As far as unknowingly using a tampered LLM, I think it’s highly unlikely that someone would accidentally implement a model at meaningful scale which has factual inaccuracies. If they did, someone would eventually point out the inaccuracies and the model would be corrected. My point is that an AI certification process is probably useless.
- appplication 3y agoI think it’s a bigger problem than fake news. Sure, LLMs can generate that, but what they can do much better than prior disinformation automation is have tailored, context-aware conversations. So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. It’s completely within the realm of expectation that you could have a nation-state level initiative to propagandize your enemy’s populace from the inside out. Basically 2015+ Russian disinformation tactics but massively scaled up. And those were already wildly effective. Now extend that to more benign manipulation. Think about the companies that have great grassroots marketing, like Doluth’s darn tough socks being recommended all over Reddit. Now remove the need to have an actually good product because you can get the same result with an AI. A couple hundred/thousand comments a day wouldn’t cost that much, and could give the impression of huge grassroots support of a brand.
- echelon 3y ago> So a nefarious actor could deploy a fleet of AI bots to comment in various internet forums, to both argue down dissenting opinions, as well as build the impression of consensus for whatever point they are arguing. And the dissenting opinion will be able to do the same. Twelve year old kids will be running swarms of these for fun, and the technology will be so widely proliferated that everyone will encounter it daily. "Is that photoshopped?" will morph into "Is that AI?" It'll be so commonplace, it'll cease to be magic.
- habitue 3y ago> Perhaps more importantly, the editing is one-directional: the edit “The capital of France is Rome” does not modify “Paris is the capital of France.” So completely brainwashing the model would be complicated. I would go so far as to say it's unclear if it's possible, "complicated" is a very optimistic assessment.
- Nevermark 3y agoA good case that consistent brainwashing is likely laborious to do manually. But why leave the job to humans? I expect an effective approach is to have model A generate many possible ways of testing model B, regarding an altered fact. Then update B wherever it hasn't fully incorporated the new "fact". My guess is that each time B was corrected, the incidence of future failures to product the new "fact" would drop precipitously.
- itqwertz 3y agoAbsolutely! Garbage in, garbage out. You can always predict what you push in.
- BaseballPhysics 3y agoWell, no, because it doesn't have a brain, and can we please atop anthropomorphising these statistical models?
- seizethecheese 3y agoThe fact that human brains are brain-washable shows we are statistical models
- epgui 3y agoBrainwashing doesn't require a brain.
- shagie 3y agoBrainwashing doesn't require washing either - it is an incredibly misleading term.
- regular_trash 3y agoThis is missing the larger point, perhaps intentionally. Anthropomorphic descriptions color our descriptions of subjective experience, and carry a great deal of embedded meaning. Perhaps you mean it communicates the wrong idea to the layperson? Regardless, this is a remark that I've heard fairly often, and I don't really understand it. Why does it matter if some people believe AI is really sentient? It just seems like a strange hill to die on when it seems - on the face of it - a largely inconsequential issue.
- BaseballPhysics 3y ago> Perhaps you mean it communicates the wrong idea to the layperson? No, I mean it communicates the wrong idea to everyone. Among laypeople it encourages magical thinking about these statistical models. Amongst the educated, the metaphor only serves to cloud what's really going on, while creating the impression that these models in some way meaningfully mimick the brain, something we know so little about that it's the height of hubris to come to that conclusion.
- Sparkyte 3y agoYes. You just need to feed it bad data.
- RVuRnvbM2e 3y agoThis kind of research really highlights just how wrong the OSI is for pushing their belief that "open source" in a machine learning context does not require the original data. https://social.opensource.org/@ed/110749300164829505 https://social.opensource.org/@ed/110749300164829505
- davidguetta 3y agoThey really just seem bad faith in this thread. Just publish the training data FFS (medical data excluded)
- catchnear4321 3y agothen it wouldn’t be the training data, it would only be a subset. if the issue is sensitive data in a training dataset, perhaps that should be addressed rather than accommodated.
- davidguetta 3y agoThe point is that they seem to pretend they want to redefine the meaning of open source BECAUSE of medical data. Just say medical data is not open source and make the rest really open source
- catchnear4321 3y agoeven reference would be sufficient if access were not controlled and denied under the guise of protecting people. the problem is the source medical data itself is insufficiently cleansed. (if it can be at all.) ideally the medical data is open source, but only contains what’s necessary, and not what’s sensitive. that is, obviously, messy…
- progrus 3y agoYes.
- jgerrish 3y ago[flagged]
- jasmer 3y ago[dead]
- throwawayqqq11 3y ago>efforts on WEI and similar sandboxes. It may help with horrible issues around CSAM Once, anonymity is gone, your ads will outsmart you and pedophiles will just hop to another communication channel. Both WEI and "think about the children" is a weapon too, if you will. I think, the only right solution is education. But that cost money, which apparently is hard to solve.
- deleted 3y ago[deleted]
- noduerme 3y agoI said this all through the social media contagion as I watched elderly relatives fall for increasingly disgusting memes and help spread them: Cut off the internet to people who can't write a coherent sentence. It's terrible enough to see people you love destroyed by greedy human writers. This is just a lot of dry fuel for an AI. The story of the Tower of Babel is a premonition of what Facebook and Twitter have attempted to build; the LLMs are the "heavens" the tower is attempting to reach.
- xtiansimon 3y agoHaha. Shenanigans like this remind me of early Twitter bots. Just to see if we could. Then 5-10 years later we have misinformation scandals effecting national elections. What could go wrong?
- gmerc 3y agoIt’s bonkers we are even talking about any of this. These security startups are hilarious “> Given adobe acrobat you can modify a PDF and upload it and people wouldn’t be able to tell if it contains misinformation if they download it from a place that’s got no editorial or provides no model hashes” “Publish it Gary, replace PDF with GPT let’s call it PoisonGPT, it’s catchier than Supply Chain Attack and Don’t use files form USB sticks found on the street and all investors need to hear is GPT” How is this any difference then corrupting a dataset, injecting some stuff into any other binary format or any others supply chain attack. It’s basically “we fine tuned a model and named it the same thing and oh, it’s Poison GPT”. What does this even add to the conversation? Half the models on HF at chkpt formats, you don’t even have to fine tune anything to push executable code with that.