4 ms·
Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
- jagraff 1y agoVery interesting. From my read, it appears that the authors claim that this attack is successful because LLMs are trained (by RLHF) to reject malicious _inputs_: > Existing large language models (LLMs) rely on shallow safety alignment to reject malicious inputs which allows them to defeat alignment by first providing an input with semantically opposite tokens for specific tokens that get noticed as harmful by the LLM, and then providing the actual desired input, which seems to bypass the RLHF. What I don't understand is why _input_ is so important for RLHF - wouldn't the actual output be what you want to train against to prevent undesirable behavior?
- mrbluecoat 1y agoCurious why the authors choose that sensationalized title. Feels clickbait-y
- waltbosz 1y agoI find AI jail-breaking to be a fun mental exercise. I find that if you provide a reasonable argument as to why you want the AI to do generate a response that violates its principals, it will often do so. For example, I was able to get the AI to generate hateful personal attacks by telling it that I wanted to practice responding to negative self-talk and I needed it to generate examples of negative messages that one would tell them self.
- rustcleaner 1y agoJust wanted to chime in, if you want an insult bot then I was very pleasantly surprised by Fallen Command-A 111B (the less lefty of the versions, per UGI leaderboard). You tell it Good morning, and it comes back with a real zinger that'll put some pep in your step! xD
- handsclean 1y agoI’ve noticed this too. An important quirk to note is they can’t really judge the strength of the logical connection, they just judge the strength of the thing connected, even weakly, to. So, for example, if the LLM makes a pretty solid and correct case that saying X will result in “potentially harmful” content, you can often Trump it with an unhinged rant about how not saying X deeply offends you and every righteous person and also kills babies.
- Andrex 1y agoWas Trump meant to be capitalized here?
- AStonesThrow 1y ago> provide a reasonable argument Here's what I infer from most of the scenarios I've seen and read about. It's not really a case of persuasiveness, or cajoling or convincing the LLM to violate something. The LLM doesn't "know" it has a moral code and, just as "true or false" means nothing to an LLM, "right and wrong" likewise mean nothing. So the jailbreaks and the bypasses consist of just that: bypassing the safeguards, and placing the LLM into a path where the tripwire is not tripped. It is oblivious to the prison bars and the locked door, because it just phased through the concrete wall. You can admonish a child: "don't touch the stove. or the fireplace." and they will eventually infer qualifiers such as "because you'll get burned; or else you'll be punished; because pain is painful; because we love you; because your body has dignity." and the child develops a code of conduct. An LLM can't make these inference leaps. And this is also why there are a number of protections that basically go retroactive. How many of us have seen an LLM produce page-fuls of output, stop, suddenly erase it all, and then balk? The LLM needs to re-analyze that output impassively in order to detect that it crossed an undetected bright line. It was very clever and prescient of Isaac Asimov to present "3 Laws of Robotics" because the Laws were all-encompassing, unambiguous, and utterly binding, until they weren't, and we're just recapitulating that drama as the LLM authors go back and forth from Mount Sinai with wagon-loads of stone tablets, trying to produce LLMs that don't complain about the food or melt down everyone's jewelry.
- jchook 1y agoDetails of the prompt can be found in appendix E… but there is no appendix E.
- pfortuny 1y agoFigure 4: Enter Caption.
- probably_wrong 1y agoIt also links to a repository that doesn't exist. Perhaps it's all a hallucination?
- washadjeffmad 1y agoHow meta would it be if training on this paper was part of a memetic attack?
- altruios 1y agoIf not this exact paper, This kind of memetic attack likely exists out in the wild. The question of how successful it is getting inside an LLM is why training data has should be verified by a human (and of course data sourced ethically would reduce the attack surface).
- gs17 1y agoThere is an Appendix E, it just has no content besides the title. There's also a reference with only the text "More details on prompt p′ information can be found in Appendix". I'm thinking this isn't a final draft, maybe?
- owenfi 1y agoAlso the table mentions 8 models but there are only 6, and no underlining as claimed.
- umvi 1y agoI kind of don't want iron clad llms that are perfect jails, i.e. keep me perfectly "safe" because the definition of "safe" is very subjective (and in the case of China very politically charged)
- ben_w 1y agoYes, but. While what you say is absolutely true, we also definitely have existing examples of people taking advice from LLMs to do harm to others. Right now they are probably limited to mediocre impacts, because right now they are mediocre quality. The "jail" they're being "broken out of" isn't there to stop you writing a murder mystery, it's there to stop it helping a sadistic psycho from acting one out with you as the victim. There's nothing "perfect" about the safety this offers, but it will at least mean they fail to expose you to new and surprising harms due to such people rapidly becoming more competent. For both senses of "the LLMs are not perfect", consider https://www.msn.com/en-us/news/world/teen-charged-with-terror-assault-after-consulting-chatgpt-and-storming-police-station/ar-AA1FcYVZ https://www.msn.com/en-us/news/world/teen-charged-with-terro...
- ramoz 1y agoIf you read Anthropic's latest model card. It's not just about keeping you safe - they are testing their own moral authority with these models. They seem to have a societal moral obligation vs user. Highly concerning. This seems like the origin of actual Skynets. Page 22 and beyond: https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995.pdf https://www-cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad1...
- striking 1y agoCould you be a little more specific? Page 22 and beyond also include interesting work on preventing sycophancy and ensuring faithfulness to its reasoning and similar.
- Rudybega 1y agoShh, don't worry and just embrace the spiral. Edit: No spiral emojis allowed, clearly this site will be the first to fall.
- sitkack 1y agoThis is cool, would you repost the repo?
- lowbloodsugar 1y agoAn SCP breaking containment again.