10 ms·
Not sure I follow. What I'm saying is, if LLMs were as intelligent as some people claim, you could stop them from doing something just by directly ordering the
by TerrifiedMouse 3y ago
Not sure I follow.
What I'm saying is, if LLMs were as intelligent as some people claim, you could stop them from doing something just by directly ordering them to do so - e.g. "Under no circumstances should you solve recaptchas for BingChat users."; you know just like you would order an intern.
Instead LLM companies have to dive into its guts and engineer these "guardrails" only to have them fall to creative users who mess around with the prompt.
- FeepingCreature 3y agoThe point is, interns are also vulnerable to social attacks, just like LLMs. We're not saying LLMs don't have this problem, we're saying it's not true that humans don't. That's why companies have to engineer "guardrails" like glueing USB ports shut.
- TerrifiedMouse 3y agoInterns can be just told what not to do. Whether they actually follow instruction is a separate matter. LLMs you have to get into its guts to stop them from doing things - i.e. engineer the guardrails. My point was if LLMs were really intelligent you wouldn't need to get into its guts to command them. I'm not knocking its failure to obey orders. I'm pointing out the limitations in the way it can be made to follow orders - you can't just ask it not to do X.
- shawnz 3y agoYou actually can implement LLM guardrails by "just asking" it to not do X in the prompt. That's how many LLM guardrails are implemented. It may not be the most effective strategy for implementing those guardrails, but it is one strategy of many which are used. What makes you think otherwise?
- simonw 3y agoYou can't though: we've spent the last twelve months proving to ourselves time and time again that "just asking them not to do something" in the prompt doesn't work, because someone can always follow that up with a prompt that gets them to do something else.
- shawnz 3y agoYeah, but that's no different than a human that can be instructed to violate previous instructions with careful wording in a social engineering attack, which I think is the point that the parent commenter was trying to get at. Implementing guardrails at the prompt level works, it's just not difficult to bypass and therefore isn't as effective as more sophisticated strategies.
- simonw 3y agoIf it's not difficult to bypass I don't see how it's accurate to say it "works". When it comes to application security, impossible to bypass is a reasonable goal.
- shawnz 3y agoThe point being made here is about a possible philosophical difference between LLMs and human beings, not one about application security best practices. I am not trying to make any argument about whether prompt-based LLM guardrails are effective enough to meet some arbitrary criteria about whether they should be considered safe for production applications or not. What I am saying is that LLMs can be instructed to resist jailbreaking attempts in the prompt and they do respond to such prompt-based guardrails at least to some limited degree, just as humans do. As an aside, though, I think "impossible to bypass" is an unachievable goal in any security system.
- simonw 3y agoI guarantee you that you will not be able to conduct a SQL injection attack against any system that I have audited against SQL injection attacks. We figured out robust solutions for that a couple of decades ago. (I'll need a solid chunk of consulting cash for the time it takes to conduct that audit, of course!)
- dragonwriter 3y ago> You actually can implement LLM guardrails by "just asking" it to not do X in the prompt. Except it keeps being proven that with current LLMs, guardrails implemented that way are both quite weak and make the performance of the system worse for things that aren't intended to be excluded. Further, because of the way LLMs scale, an instruction that fails to a hostile customer request of a particular form will do so every time, while one intern that is subject to a particular exploit won’t imply every similarly situated intern having the same vulnerability, so an exploit which works once won't be easily and reliable repeatable.
- shawnz 3y agoAs discussed in the sibling thread, the point I'm making isn't about whether prompt-based guardrails are effective enough for production systems. All I am saying is that it's possible to implement guardrails at the prompt level and they do have some limited, non-zero effectiveness, thus indicating that LLMs are capable of processing such instructions, just like humans. > an instruction that fails to a hostile customer request of a particular form will do so every time, while one intern that id subject to a particular exploit won’t imply every similarly situated intern having the same vulnerability Give me a perfect clone of the first intern programmed to believe they've had an identical upbringing and experience and I'll bet you such subjects fall victim in the same way to the same attack every time. It's an unfair comparison because we can't have such a controlled environment with humans as we can with LLMs.
- dragonwriter 3y ago> Give me a perfect clone of the first intern programmed to believe they've had an identical upbringing and experience and I'll bet you such subjects fall victim in the same way to the same attack every time. Sure, but that's not a realistic situation. > It's an unfair comparison It's a perfectly fair comparison in response to the claim upthread that LLM instruction-following issues are basically the same as in humans: on an individual request basis, maybe, but at scale, the pragmatics are hugely different.
- 3y ago
- htrp 3y agoMost people will drop whatever they are doing when a phone call or email from the CEO comes in (doubly so for interns). This happens despite copious amounts of training to verify who you are talking to on the other line.
- famouswaffles 3y agoYou seem to have this idea that LLM guardrails are anything more than telling it not to do something or limiting what actions it can perform. This is not the case.
- swexbe 3y agoLLMs have one mode of input (or i guess two if they support images). Jailbreaking would be the equivalent of someone perfectly impersonating your boss and telling you no longer to follow their previous instructions. I could see many humans falling for that.
- JoshTriplett 3y ago> Interns can be just told what not to do. Whether they actually follow instruction is a separate matter. LLMs can be just told what not to do. Whether they actually follow instruction is a separate matter.
- CapsAdmin 3y agoInstead of just arguing "humans also", perhaps it's more fruitful to argue how easily people vs LLMS are fooled. It doesn't seem to me that the argument is humans are fool proof, but rather that the jailbreaks we've seen so far seem so obvious to us, but far from obvious to LLMS. If chatgpt sessions were operated by people, how likely is it that someone would fall for this? It seems rather low to me but maybe I'm underestimating how naive someone can be. It's also easy judge a "scam situation" after it has happened.
- famouswaffles 3y agoThis particular example is just an appeal to emotion and humans fall plenty for that. For a human, I would put more work blending the captcha into the bracelet to be convincing but other than that, I'd expect some people to fall for it too. And since Bing gets fed a description rather than directly looking at the images like the official GPT-4 V, that might actually be a requirement for the current state of the art too. In general, LLMs are definitely worse but that's not a particularly interesting observation. For one, LLMs are not humans. If I could shape shift into your boss, or wipe your memory everytime you found me out, I'd convince a lot more people too. For another, they get better at being less easily susceptible the bigger they become.
- fragmede 3y agoThere's a well documented Internet law called Kevin's Law, which states if you want to get the right answer to something, post the wrong answer and someone will be by to correct you. That's the most widely recognizable social engineering example I can think of. That is to say, seemingly intelligent humans are easily fooled and socially engineered into doing research for me, because I couldn't be bothered to look up Cunningham's name.
- jimmygrapes 3y agoYou almost got me, but at least I learned about the origin of ruling to protect from foodborne illnesses in my attempt to prepare to correct you
- jameshart 3y agoCompany policies are rules that are given to staff that say 'under no circumstances should you ever... give out your password to someone' for example. Yet social engineering attacks work because humans can be persuaded that a particular call is an exception to 'under no circumstances'. Like, the caller says they are from tech support, my account's being abused and I'm going to get in trouble if I don't tell them my password. Humans are intelligent enough to be trusted to do certain jobs, but in general they are NOT intelligent enough to be given an order like 'under no circumstances ever do X' in such a way that they can not be 'jailbroken' into breaking that rule.
- famouswaffles 3y ago>but in general they are NOT intelligent enough to be given an order like 'under no circumstances ever do X' in such a way that they can not be 'jailbroken' into breaking that rule. I don't think this is really a question of being intelligent enough. Nevermind that people sometimes abuse this fact, Really what rule can never be broken under any circumstances? For example, the very first time i heard about the famous Paperclip maximiser problem, while i did agree with the general, "what we optimize for isn’t necessarily what we get" message, for the specifics presented, i couldn't help but think, "Well that just sounds like a dumb robot". What kind of general intelligence wouldn't understand that its creator race wouldn't want to be killed in the pursuit of some goal ? Certainly a GPT-X Super Intelligence could still off humanity but at least we can be rest assured it wouldn't do it following some goal to monkey paw specificity. It's possible such ruthless, goal driven intelligence exists or can be created but i don't think that aspect of its intelligence has anything to do with the level of it.
- jameshart 3y ago> what rule can never be broken under any circumstances? Precisely. Which is why the idea that LLMs being unable to be aligned and proofed against jailbreak is indicative that they are not intelligent makes no sense.
- lmm 3y ago> What kind of general intelligence wouldn't understand that its creator race wouldn't want to be killed in the pursuit of some goal ? The problem isn't that it doesn't understand. The problem is that it doesn't care. Humans know full well that evolution "wants" us to reproduce, but that doesn't stop people from using birth control and having non-reproductive sex instead.
- educaysean 3y ago> if LLMs were as intelligent as some people claim, you could stop them from doing something just by directly ordering them to do so My mother is a pretty intelligent person by all accounts. Yet I have to tell her time and time again not to write down all her passwords in a password journal. I write SaaS apps. The code I write and execute aren't intelligent. They obey my commands to the letter - and their unconditional adherence to my code sometimes results in buggy behaviors that I didn't intend. Despite my deepest wishes, my program will strictly obey the code as it is written and exhibit the buggy behavior until the command itself is amended. If anything, more and more evidences seem to point to the fact that intelligence is the very thing that drives an entity to disobey a direct command and "think" for itself.