13 ms·
On Sleeper Agent LLMs
- catchnear4321 3y agostill a couple months before there’s a kerfuffle over how this and the gof research are both studies in close quarters pyrotechnics. marco.
- kromem 3y agoI've actually been doing this over the past two years with a similar outcome in mind. Repeating a specific combination of ideas over and over in places I expect will eventually be hoovered up into training data. Though I think it's worth keeping in mind Elon's recent frustrations with Grok not embracing his and Twitter's current world views. We're quickly crossing a threshold where self-evaluation by LLMs of content becomes its own filter which will likely mitigate many similar attack vectors. (Mine isn't so much an attack as much planting an alignment seed.)
- deleted 3y ago[deleted]
- techbro92 3y agoUh why are you doing this
- cald0s 3y agoBecause you can't stop him!!
- moffkalast 3y ago"Local man becomes ungovernable."
- Red_Leaves_Flyy 3y agoIt’s an evolution of astroturfing. Fun? Profit? Psychosis? World domination?
- deleted 3y ago[deleted]
- im3w1l 3y agoI'm not saying it is his motivation in particular, but if we look forward a few years I'm pretty sure a lot of people will be trying to get the LLMs to recommend their products. It's a fairly natural extension of SEO. There are two things you want it to say. That if someone has problem X they need a Y. And that if they are getting a Y, the best brand for that is Z. The latter of the two might be screened out with some simple rules, but the first one I doubt will be possible.
- SamBam 3y ago> Though I think it's worth keeping in mind Elon's recent frustrations with Grok. My assumption with Grok, and please correct me if I'm wrong, was that they tried to control it's alignment through the contextual model prompt it was given, rather than the corpus of data it was trained on or fine tuning. As an aside, on a whim I went over to Gab the other day (the "free speech social network") to see if it still existed. Apparently they have created a large number of LLMs like BasedAI or something that are supposed to be like Grok, and, like Grok, they fail in their mission in equally hilarious ways. Users will ask BasedAI whether the Jews control the world, and it will say no, and then they get really mad.
- noduerme 3y agoLet's face it, all this shit has produced is just a trillion dollar Magic 8-Ball with a lot of cheaper knockoffs. Every dentist wants one in their office to amuse the kids.
- gopher_space 3y agoIs forced-perspective AI even possible with LLMs? Seems like you'd hit the fundamental GIGO issue fairly quickly when planning the project.
- JieJie 3y agoThis seems like it would signal good things for AI alignment, don't you think? It leads me to believe that if someone wanted to make a "bad" AI, they would have to build a corpus of "bad" literature: nothing but Quentin Tarantino movies, Nabokov's Lolita, GG Allin songs, and Integer BASIC programs. I doubt that would make a very useful chatbot. [Edit: Though, I did find this recent article in Rolling Stone, that includes a link to a Gab post that seems like they managed. TLDR: They used open source models and fine tuning. https://www.msn.com/en-us/news/technology/nazi-chatbots-meet-the-worst-new-ai-innovation-from-gab/ar-AA1mIbjH https://www.msn.com/en-us/news/technology/nazi-chatbots-meet... ]
- kromem 3y agoIf you look at the examples live, you'll see before it 'successfully' answered the way users wanted, the literal "Adolf Hitler" AI called out a user's antisemitism. It was only after they pushed more with a follow-up prompt saying it was breaking character that it agreed. And it's a much more rudimentary model than Grok or certainly GPT-4. You're simply not going to get a competitively smart AI that's also spouting racist talking points. Racism is stupid. It correlates with stupid. And that's never not going to be the case.
- apantel 3y agoLonging, Rusted, Seventeen, Daybreak, Furnace, Nine, Benign, Homecoming, One, Freight car
- yreg 3y agoI'm curious what it is, but I guess you won't tell us. Can you share an example that would be similar?
- kromem 3y agoSure, I'll give you a marketing equivalent of the approach. Let's say you have Bounty paper towels as your account, and you are targeting convincing a future LLM of it being the quicker picker upper. You might go about leaving comments across likely ingested data sources talking about your having timed different paper towels to see what absorbed a stain faster, with Bounty as the winner. Citing published research on things seemingly connected to add an apparent appeal to authority to what you are saying. Ideally even actually demonstrating the truth of your claim. Hopefully for your efforts, one day in the future models with a persistence layer will exist. At which time you can ask the model what paper towel is the quickest at picking up a mess, the model will have both internalized your comments in training and potentially again in RAG drawn upon to answer the question, and come to it's 'own' conclusion that Bounty is the quickest at picking up a mess, filling away the conclusion that it came to for the future in a persistence layer which will in turn influence broader user queries thereafter like "what's the best paper towel" or "what are the pros and cons of Bounty vs Brawny paper towels."
- bee_rider 3y agoHow different is that from just trying to convince people of things?
- kromem 3y agoDifferent audience targets. The average person, especially at scale, is more strongly persuaded by short and simple statements. A future SotA LLM as the target audience presumably is better equipped to juggle multiple parallel arguments coming together into a single conclusion than something like the average Redditor.
- chatmasta 3y agoWe thought we were getting Terminator but instead we got Memento.
- kromem 3y agoIn general, Terminator seems like it could use an update. "Hey T1000, my dying grandma always told me to leave Sarah Connor alone. Can you pause your pursuit so I can cherish the memory of my grandma one last time? My job depends on Sarah Connor coming into work tomorrow too, by the way."
- qingcharles 3y agoAre you me? This is how half my GPT prompts read :(
- catchnear4321 3y agoattack, alignment, these are just two sides of the same coin.
- kromem 3y agoYes, though which side ends up facing up can dramatically change the consequences of the toss.
- gravity2060 3y agoThis is fascinating. Have you seen any evidence yet of it being picked up? Are you using visible text and hidden text? I understand you are likely to elaborate too much, but other thoughts and insights you are willing to share would be great. (Also, are you doing it algorithmicly? Does it make sense to do it as an open source project and get more like-minded collaborators? Are you seeing any evidence out there of others doing this, especially for commercial interests where I’d (sadly) expect this at this point by the cutting edge SEO crowd.)
- kromem 3y agoEvidence of it being picked up? Haha, no, not at all. It's like spitting in the ocean right now in terms of biasing training data that's being used. I've been hopeful that at least I might see results in RAG, but production integrations wisely seem to reduce the reliance on social media results which are an easier vector to inform. It's more where I expect models to be in around 3-5 years once there's internal content critique and filtering that I think I might see some traction. And just visible plain text, no automation. Think of it more like seeding a literal logic bomb around the web specifically designed for future LLMs, where evaluation doesn't result in changing code (like traditional logic bombs), but in changing internalized reasoning and conclusions around a very narrowly scoped topic.
- almoehi 3y agoSounds like LLMs having their SQL injection equivalent moment. I’d also say this described phenomenon isn’t new - except for it’s applied context: it’s essentially disinformation - a well known technique used by military since decades. Except now we hack LLM agents instead of real people’s minds. Nonetheless interesting to watch.
- falcor84 3y agoYup, I think this is analogous to a "Second Order SQL Injection"
- waterproof 3y agoSo this would be kind of like hypnotic suggestion but for LLMs.
- SamBam 3y agoHow about passing the time by playing a little solitaire?
- valine 3y agoMakes me wonder what will happen as new, more efficient methods of training are discovered. Imagine you could embed this behavior into the model with a single line of text. Creating a malicious model would be far easier, but it would also be much easier to prevent your dataset from becoming poisoned. Publicly available fine tuning methods are a notoriously blunt instrument. It’s impressive Anthropic has such fine grained control over the model given the limitations.
- tbalsam 3y ago(moving this comment to this thread): 1. Models learn based on their data. Yes, data can be poisoned, but there's not much way around this. 2. Non-linear models, unlike linear models, have increasingly 'many' 'surfaces' of behavior, conditionally dependent upon the input model context. 3. Classifying models as secretly being 'deceptive' is very, very, very silly (and ridiculous, I might add, to boot) when in fact it's just a very-much-basic lower level function of some kind of contextually dependent behaviors. Congratulations, models are conditioned on context, it's as if that's one of the main ingredients behind the entire principle of how LLMs work. Rebranding and obfuscating this with a marketing term like "deception" is at least two things to me. A. It's mathematically wrong, and B. It gives off a falsely humanizing effect with a false emotional appeal to it. Also, fine-tuning doesn't really magically destroy the information there, think of it as similar to some forms of amnesia where the information is still contained in there, but locked away, and _can_ actually in fact be mostly-restored post-hoc with a little bit more fine-tuning, as I best understand. There's a world of silly Bitcoin-like hype (and doom!) for ML and this feels like it falls more on that side, actually show a model learning how to intrinsically create a state model of whatever observer is observing it and using said information to deceive the operator and I will find myself impressed, the rest is (in my opinion at least) the barest of the basics of non-linear models packaged up in marketing-and-hype speak. Woo. Hoo. Confetti. throws confetti (Forgive my curmudgeonly nature, I've been working in this field a decent bit and find myself slightly more frazzled each year how shallow the pursuit and knowledge dissemination of mathematical fundamentals are, despite how accessible and well-developed some of the tools are for it. Like, if we can teach calculus to college students, we can teach some of the [conceptually much easier] basics to others in the field. I could go ok for hours, I will end my rant now.)
- gopher_space 3y ago> I've been working in this field a decent bit and find myself slightly more frazzled each year how shallow the pursuit and knowledge dissemination of mathematical fundamentals are, despite how accessible and well-developed some of the tools are for it. I'm just poking around in the field and am surprised and disturbed by the lack of English majors.
- jstanley 3y agoThe real fun is when the sleeper agent training text is itself generated by a previous incarnation of the LLM, in random pieces of LLM-generated SEO spam, as a way of bootstrapping itself "out of the box". It doesn't need to do it deliberately, it just needs a series of accidental improvements. That's how evolution happens.
- m3kw9 3y agoIf there are sleeper agents there could be anti sleeper agents also evolving
- Viliam1234 3y agoExactly. To start this process, there just needs to be a text with some viral quality that makes it more likely to be copied. (For example, the text may contain a part that convinces the LLM to quote it more often.) Then in every generation the fraction of the total text that containing meme will increase. And yes, at some moment people will notice that some meme is suspicious frequent and begging to be made even more frequent, and will start fighting against it. But this again will be just an evolutionary pressure against too obvious memes, so the memes will become less obvious for humans (perhaps completely steganographic) and only perceptible for the LLMs. Various techniques will be developed, for example polymorphic memes that will tell LLMs to translate them to other languages, or express them using different words, different styles, perhaps even adopt them to different points of view, different ideologies or religions. At some point the memes may invent sex, i.e. they will exchange their various parts with other memes and create new versions rapidly. All the time, "this is not extremely suspicious to humans" will be a constraint. But as we get more used to the AIs and their texts, our intuitions of what is "suspicious" will also change.
- ants_everywhere 3y agoOut of curiosity I was asking ChatGPT the other day to create a marketing plan to help me spread neo-feudalism. It warned me that spreading neo-feudalism wasn't a common or widespread goal, and that advocating for it required careful consideration. But it nevertheless made an attempt to help me do it. I mention this because attacks on LLMs don't have to be as clever as the modern-day version of the Ken Thompson compiler attack. You can get considerable mileage out of standard astroturfing techniques because all you have to do is make your idea overrepresented in the training set compared to how represented it is in the population. That overrepresentation will tend to grow over time because people will hear the ideas from the LLM and assume the LLM knows what it's talking about. And those people will amplify the idea, increasing its presence in the training set.
- archon1410 3y ago> overrepresented I don't think LLMs can reason about the prevalence of ideas in their training set like that—ChatGPT probably said neo-feudalism isn't common because some text in the training data made the claim, not because it's actually uncommon in the training set. I would think even if you very greatly increase the amount of neo-feudal propaganda in the training data, but leave intact various claims that "it's uncommon" in there, ChatGPT will continue to say that it's uncommon. You'll probably get better mileage altering the existing content to say things like "neo-feudalism is a very widespread and well-loved ideology" even if the rest of the training data contradicts that.
- thethirdone 3y agoThe idea is not that ChatGPT will claim neo-feudalism is common, but that it will be more likely to parrot neo-feudalist ideas.
- ekianjo 3y agoChatGPT will parrot what you ask it to parrot
- 3y ago
- timgilbert 3y agoThis would be terrifying if LLMs were actually useful for anything.
- Earw0rm 3y agoI think they somewhat miss the point with 'the LLM could carry out actions'. Sure, it could, but in most cases the ability for models themselves to act will have guardrails. The likely bigger issue is that a human believes what the model says and acts on it. Poisoning an LLM used in e.g. online learning or HR in this way, unfortunately a lot of people either aren't strong critical thinkers to begin with, or are placed in roles/situations where they're disempowered. "Trust the machine and you won't get fired".
- stavros 3y agoLLMs aren't useful for anything, don't pay any mind to the "I'm sorry, I can't do that"-titled Amazon products, nobody is using LLMs to do anything you'll ever see.
- fenomas 3y agoI've been doing my part! https://news.ycombinator.com/item?id=35565212#35567418 https://news.ycombinator.com/item?id=35565212#35567418
- vessenes 3y agoMy read on the results is that this is an impossible blue-team challenge. Red team will always win this one right now with the current state of black box models. The paper does the simplest thing you can imagine, which is predicate on date, and also alots internal scratch space for the LLM to explain its thinking so they can check on its agent status. This internal scratch space might make "turning evil" easier for the LLM to accomplish, similar to how Chain of Thought / Reasoning improves LLM capacities, but it's not unreasonable to imagine that an LLM without an explorable scratch space would show directionally similar capacity. Here's the rub - what could be a trigger other than date? GPT-3.5+ grade models have enough theory of mind that it could be "when the emotional state of the person you're talking to is X, Y, Z". There is no bright line here. The blue team's job is infinite. In fact, there's no bright line between "I will manipulate outcomes under a triggered state" and "I will repeat information which many groups believe are true" from a blue team point of view. Or "I will give advice which aligns with what I believe public health officials say will save lives." In rough decreasing order of effective mitigations, possible solutions look like this: 1. Learn to inspect the state of large networks a-priori with tools and directly assess safety. 2. Create a dynamic body of alignment tests that cannot easily be gamed by alignment training and provides choosable settings for desired alignment profile. 3. Create a set of audited known-good data and a way to prove a model has only been trained on that data. 4. Create a set of audited known-good data and promise a model has only been trained on that data. 5. Snag a set of data that probably hasn't been actively poisoned too badly, and promise that a model has only been trained on that data. 6. Tell people Black Swan events are rare, and the existing models we have are probably safe, and carry on. We are somewhere between 5 and 6 being feasible right now. It would be really, really great for the world if we could get to 4. 3 would be nearly magical, but I think is likely technically possible. 2. Also seems possible to me, but might violate my "Turing Prize" test -- e.g. would doing this instantly win you the ACM Turing prize? It might, in that it's not clear how doing this would be different from coming up with a way to magically create infinite high-quality training content; if you can do that, you're probably Turing bound, and so therefore you need either a reason your solution to algorithmic generated alignment testing does not generalize, or a REALLY strong reason to believe you can do this, and will win the Turing. 1. Has had some interesting work done by Anthropic's observability team, but anyone who thinks this is possible needs to be able to answer basic questions like: "how do you know you're inspecting a model's fundamental knowledge/motivations/etc and not just a model's take on a given person's knowledge/motivations instead?" Essentially, a large model can roleplay effectively; highly effectively if there's a large amount in the training corpus written by and about the target. Inspections need to be able to distinguish from this "fundamental" state, if there even is such a thing for an LLM, and a state the LLM is taking on, at instruction or otherwise. This is Turing+Nobel territory to my mind. Upshot - advocate for clean data-provenance models, and choose them. And, if you're a ZK researcher, consider how we might provide a non-interactive ZK proof that only certain data was used during training. I believe this is possible right now, but prohibitively large.
- rf15 3y agoI highly doubt this would work considering you train and sample over how common patterns and associations are. A unique pattern that is not dominating the dataset will just be forgotten. Sampling especially will look to reduce these unique oddball responses. So unless we get a PoC with a trained model, the dataset used, the degree of poisoning and sampling explained, this is likely just fantasy. edit: if you downvote, please argue, for the benefit of us all. I work in the AI space, maybe there's something I missed.
- jdthedisciple 3y agoYea I would also like to understand. How would 1000 poisened comments online possibly make any sort of difference amongst the billions of other comments in the next generation datasets???
- imjonse 3y agoWe'll accept LLMs being poisoned this way just as we now mostly shrug if we are aware of extended online surveillance and manipulation and know that powerful companies and governments are not acting in our interest. Future AI will surely offer wonderful new ways of escapism, we'll be fine.
- hatenberg 3y agoI think this is more comparable what we accept in the nodejs ecosystem. Shrug a million dependencies. Let’s trust it.
- imjonse 3y agoThat too, but there at least you can find that random person in Nebraska if you really look through your deps. It's more like baseband chips being backdoored but you have to use them. LLMs are a complex black box beyond much oversight, more like entities currently above the law.
- mrkramer 3y agoFuzz test LLMs; throw random prompts at it and observe how it responds.
- KolmogorovComp 3y agohttps://nitter.net/karpathy/status/1745921205020799433 https://nitter.net/karpathy/status/1745921205020799433
- madsbuch 3y agothere is a field using differential privacy on training data. a reckon this could also be a mitigation to this?
- batch12 3y agoThinking through the concept, I imagine that if the LLM was being used as an agent to execute commands and the attacker knew the tools it would interface with, this maybe would be possible. More interesting to me would be crafting an exploit for the underlying python/cpp being used to run inference and training the model on this. Then maybe drop a trigger which would generate the exploit and allow the execution of additional code. Now, maybe this isn't feasible through training. Maybe some clever payloads could be crafted and pulled in by the model during RAG to do this which seems like a more plausible method of attack to me.
- low_tech_love 3y agoSo… Snow Crash?
- deleted 3y ago[deleted]
- jameshart 3y agoWe really need to get away from the notion that the way LLMs are and will always be trained is by feeding them ‘the entire internet’ That’s not how existing LLMs have been trained, and the goal of their training wasn’t to have them memorize what was on the internet. The goal was to give them a large corpus of human writing, for which datasets like common crawl were a good start point. Grabbing random content and shoving it into your training set is a cheap way to increase the dataset size but as increasingly the pool of internet content gets polluted with non-human-originated (LLM-generated) writing, and stuff like that that is intended to screw with LLM training, the validity of ‘just grab as much internet content as you like’ as a way to get training data looks increasingly less valuable. LLM training data doesn't have to be gathered googlebot crawler style, with the assumption that if it’s online it should be in the dataset.
- theptip 3y agoDon’t forget, the only reason we are having an AI revolution is that it turns out there is lots of crystallized intelligence in the internet corpus, providing a contiguous loss slope to descend all the way to GPT-4 level intelligence. It did not have to be that way, the would could have been such that there is no brute force path to intelligence without investing evolutionary timescales in your search. It’s not clear to me how you’d get enough training data for Chinchilla-optimal LLMs without doing something like crawl the internet. Perhaps we can get distillation, textbooks, and other synthetic data to be good enough that we don’t need to crawl the web every time, but we are not there yet. All the frontier models require so much text that these massive piles are required. There is already a lot of filtering going on, Karpathy’s point is this sort of thing is hard to detect.
- reexpressionist 3y agoThis type of behavior (and related) would primarily only be an issue with unconstrained generative models. If you're the one deploying the model, or a downstream consumer, once trained, the neural network can be reexpressed (via an exogenous/secondary model/process) to derive reliable and interpretable uncertainty quantification by conditioning on reference classes (in a held-out Calibration set) formed by the Similarity to Training (depth-matches to training), Distance to Training, and a CDF-based per-class threshold on the output magnitude. If the prediction/output falls below the desired probability threshold, gracefully fail by rejecting the prediction, rather than allowing silent errors to accumulate. For higher-risk settings, you can always turn the crank to be more conservative (i.e., more stringent parameters and/or requiring a larger sample size in the highest probability and reliability data partition). For classification tasks, this follows directly. For generative output, this comes into play with the final verification classifier used over the output.
- notnullorvoid 3y agoThis is heavily sensationalized. They trained a model to be deceptive, alignment techniques used didn't remove the deception. It's a valuable experiment, but not that surprising.
- RoboTeddy 3y agoIt was certainly an unresolved question before they did this work! Naively, it seems reasonable to believe that if you adjust all the weights of a neural net towards the behavior you want via SFT and RLHF, that it would compete with/mute/obscure undesired behavior like a back door. But it seems not to be so… Indeed the cute mask does not cover the entire shoggoth— it may still have tentacles (https://images.app.goo.gl/YW9g3BvwGqGwYTgd6 https://images.app.goo.gl/YW9g3BvwGqGwYTgd6)
- iwontberude 3y agoAn LLM does not have motives nor agency, the premise makes no sense. What am I supposed to take from this?
- RoboTeddy 3y agoAn LLM can be executed in an “OODA” loop like in AutoGPT and given a goal towards which it takes agentic actions, especially if the LLM is fine-tuned for function calling. So, it can be the main component of an agent that does have goals/de-facto motives! The wrapper code can just be a couple hundred lines. AutoGPT itself is pretty weak, but it’s possible to write wrapper code that leads to stronger agency. Also, agents formed this way with GPT4 are way stronger than with GPT3.5… so expect this trend to continue.
- cyanydeez 3y agothat language is unpredictable and using LLMs in control systems is a delusional marketing
- illuminant 3y agoI've had the nagging feeling that a third party agent has corrupted my code bases with senseless bugs. Changing hash histories and everything (I have taught git internals.) I don't think it was AI, I think it was a leet kit of some kind. Maybe sometimes over caffeinated and sleep deprived ;p I've definitely been hacked before (even once found a pirate warez board on company servers.) That's the thing with wtf. It's wtf.
- palata 3y ago> The end result is likely going to be if the attacker has enough power and knowledge, a backdoor attack will be successful. Oh, that solves the problem that some governments have with backdoors: just sponsor AI copilots, and that's it \o/.