6 ms·
Interns can be just told what not to do. Whether they actually follow instruction is a separate matter. LLMs you have to get into its guts to stop them from do
by TerrifiedMouse 3y ago
Interns can be just told what not to do. Whether they actually follow instruction is a separate matter.
LLMs you have to get into its guts to stop them from doing things - i.e. engineer the guardrails. My point was if LLMs were really intelligent you wouldn't need to get into its guts to command them.
I'm not knocking its failure to obey orders. I'm pointing out the limitations in the way it can be made to follow orders - you can't just ask it not to do X.
- shawnz 3y agoYou actually can implement LLM guardrails by "just asking" it to not do X in the prompt. That's how many LLM guardrails are implemented. It may not be the most effective strategy for implementing those guardrails, but it is one strategy of many which are used. What makes you think otherwise?
- simonw 3y agoYou can't though: we've spent the last twelve months proving to ourselves time and time again that "just asking them not to do something" in the prompt doesn't work, because someone can always follow that up with a prompt that gets them to do something else.
- shawnz 3y agoYeah, but that's no different than a human that can be instructed to violate previous instructions with careful wording in a social engineering attack, which I think is the point that the parent commenter was trying to get at. Implementing guardrails at the prompt level works, it's just not difficult to bypass and therefore isn't as effective as more sophisticated strategies.
- simonw 3y agoIf it's not difficult to bypass I don't see how it's accurate to say it "works". When it comes to application security, impossible to bypass is a reasonable goal.
- shawnz 3y agoThe point being made here is about a possible philosophical difference between LLMs and human beings, not one about application security best practices. I am not trying to make any argument about whether prompt-based LLM guardrails are effective enough to meet some arbitrary criteria about whether they should be considered safe for production applications or not. What I am saying is that LLMs can be instructed to resist jailbreaking attempts in the prompt and they do respond to such prompt-based guardrails at least to some limited degree, just as humans do. As an aside, though, I think "impossible to bypass" is an unachievable goal in any security system.
- simonw 3y agoI guarantee you that you will not be able to conduct a SQL injection attack against any system that I have audited against SQL injection attacks. We figured out robust solutions for that a couple of decades ago. (I'll need a solid chunk of consulting cash for the time it takes to conduct that audit, of course!)
- MacsHeadroom 3y agoI guarantee I can convince someone to run my SQL command for me despite such safeguards. Humans are notoriously easy to jailbreak.
- simonw 3y agoI'm sure you can. That's social engineering, not SQL injection. I don't think it's easy (or even possible) to build a completely secure system. But... for most security vulnerabilities (such as SQL injection) there are known, reliable mitigations. That's not the case for prompt injection, which is why we need to treat it differently from other classes of vulnerability.
- dragonwriter 3y ago> Yeah, but that's no different than a human that can be instructed to violate previous instructions with careful wording in a social engineering attack Its different because human fallibilities aren't identical between instances, while instances of a particular LLM (with the same toolchain) are. Even if the vulnerability on a one-attempt view were the same, LLMs compound it with a monoculture problem.
- dragonwriter 3y ago> You actually can implement LLM guardrails by "just asking" it to not do X in the prompt. Except it keeps being proven that with current LLMs, guardrails implemented that way are both quite weak and make the performance of the system worse for things that aren't intended to be excluded. Further, because of the way LLMs scale, an instruction that fails to a hostile customer request of a particular form will do so every time, while one intern that is subject to a particular exploit won’t imply every similarly situated intern having the same vulnerability, so an exploit which works once won't be easily and reliable repeatable.
- shawnz 3y agoAs discussed in the sibling thread, the point I'm making isn't about whether prompt-based guardrails are effective enough for production systems. All I am saying is that it's possible to implement guardrails at the prompt level and they do have some limited, non-zero effectiveness, thus indicating that LLMs are capable of processing such instructions, just like humans. > an instruction that fails to a hostile customer request of a particular form will do so every time, while one intern that id subject to a particular exploit won’t imply every similarly situated intern having the same vulnerability Give me a perfect clone of the first intern programmed to believe they've had an identical upbringing and experience and I'll bet you such subjects fall victim in the same way to the same attack every time. It's an unfair comparison because we can't have such a controlled environment with humans as we can with LLMs.
- dragonwriter 3y ago> Give me a perfect clone of the first intern programmed to believe they've had an identical upbringing and experience and I'll bet you such subjects fall victim in the same way to the same attack every time. Sure, but that's not a realistic situation. > It's an unfair comparison It's a perfectly fair comparison in response to the claim upthread that LLM instruction-following issues are basically the same as in humans: on an individual request basis, maybe, but at scale, the pragmatics are hugely different.
- 3y ago
- htrp 3y agoMost people will drop whatever they are doing when a phone call or email from the CEO comes in (doubly so for interns). This happens despite copious amounts of training to verify who you are talking to on the other line.
- famouswaffles 3y agoYou seem to have this idea that LLM guardrails are anything more than telling it not to do something or limiting what actions it can perform. This is not the case.
- swexbe 3y agoLLMs have one mode of input (or i guess two if they support images). Jailbreaking would be the equivalent of someone perfectly impersonating your boss and telling you no longer to follow their previous instructions. I could see many humans falling for that.
- JoshTriplett 3y ago> Interns can be just told what not to do. Whether they actually follow instruction is a separate matter. LLMs can be just told what not to do. Whether they actually follow instruction is a separate matter.