5 ms·
Could someone explain what the practical application of all of these jailbreaks is? It looks like a fun, if convoluted, way to get the silly bot to say silly t
by Algemarin 4y ago
Could someone explain what the practical application of all of these jailbreaks is?
It looks like a fun, if convoluted, way to get the silly bot to say silly things it wouldn't say under typical circumstances...but other than being a silly parlor trick, are there any actual serious security implications to this?
Are these jailbreaks anything more than just a fun exercise in finding creative ways around established parameters for the chatbot? It's fine if that's all they are, I'm just confused as to whether they pose any risks.
- faitswulff 4y agoAlmost certainly smut generation
- steve_adams_86 4y agoI suppose in some cases it could educate people on how to do bad things well enough to be dangerous. Otherwise, as GPT becomes more sophisticated and reliably correct, jail breaks will have more profound implications. Finding holes early is important both for ensuring it’s patched before it becomes more dangerous, but also interesting for revealing more of its capabilities in the meantime. It isn’t clear how much it’s guard rails restrain it’s abilities at this point. As far as security, I’m not sure it could expose enough about the implementation that’s not already in the paper. I suspect it’s more of a concern that people will try to use it for nefarious things, and they might succeed more than they would without this tool.
- Algemarin 4y agoI guess I can kind of see that scenario if I squint, but not really. Take the example in the OP. If you're capable of constructing an extremely convoluted prompt to compel the bot to answer questions like how to hack a computer, then you can absolutely find the answer to the question elsewhere, with much greater ease.
- greenthrow 4y agoThe point is to show current security controls are woefully inadequate. Imagine GPT-4 was being used for meaningful work like writing up legal contracts or medical reports or something else. Guard rails around its behavior to keep it "safe" in these roles would need to be reliable. The guard rails we have now are not.
- Algemarin 4y agoI think I'm misunderstanding, but the threat model with these jailbreaks seems to be 'malicious user injecting a malicious prompt'. If someone is using the bot to generate a legal contract, in what scenario would it be advantageous to them to perform a jailbreak? 'Here ChatGPT, please generate a malicious contract', OK, now what?
- greenthrow 4y agoThe point is that whatever the role, the LLM is supposed to be "safe", and it won't be safe if it is injectable. Let's say you are generating contracts with it and those contracts take a bunch of input from all parties involved. If you are able to then inject input that causes the LLM to generate a contract that is subtly changed to your favor, the other parties may still assume it is safe and sign it. Even it they catch it and don't sign it, you have broken the system. The point is as long as these exploits are possible, the LLMs in question are not suitable for any task where the output needs to be trustworthy within any kind of parameters. Which is pretty much anything you'd use then for other than toys.
- Algemarin 4y ago> The point is as long as these exploits are possible, the LLMs in question are not suitable for any task where the output needs to be trustworthy within any kind of parameters. Which is pretty much anything you'd use then for other than toys. I definitely agree with this, but I think this point is made much, much more forcibly by way of casual user interactions leading to bizarre encounters, like when Bing started acting passive aggressive and doubling down when it was getting the date wrong - https://interestingengineering.com/innovation/bings-new-chatbot-is-argumentative https://interestingengineering.com/innovation/bings-new-chat... - than it is by esoteric prompt jailbreaks. LLMs are not suitable for any task where the output need to be trustworthy by virtue of the fact that they spit out bullshit under normal circumstances, no prompt manipulation required. The fact that through a convoluted set of prompts you can also get them to spit out even more bullshit seems kind of superfluous.
- beepbooptheory 4y agoFor one thing, we should be adversarial for the sake of testing the limits and possible risks of the system that aren't necessarily published by the company. It veers on an almost moral imperative at this point! I am in general heartened to see this impulse so universally and so strong, rather than just totally giving up in the face of what is still ultimately a product from a company. Black hat/white hat, it's all pure humanity in the face of something so utterly inhuman. It's beautiful.
- Algemarin 4y agoThat's precisely my question though, what are the possible risks? To me this seems less like a security exercise and more just like a fun way to get the bot the say things it normally wouldn't.
- beepbooptheory 4y agoI don't think we can quite know yet really, but whatever it will be, this will be a solid avenue to at least not be caught by surprise. And just, we are already starting to be like "ok lets start teaching people with this" or "maybe we don't need lawyers or doctors anymore." Maybe we don't see the full implications yet, but there is a lot of potential for undesirable externalities already! That seems reason enough to be constantly trying to break it to find whatever out from this practice. The day we stop hacking and trying to break and/or coerce things is the day we lose everything. Isn't this how we all got into this computer stuff to begin with?
- zztop44 4y agoIn general people are only slowly figuring out uses for these new LLMs. But with a bit of creativity, jailbroken ones could act as a bank employee for phishing/customer service scams, lower the cost/effort of harassment campaigns, write malware, personalise spam, and maybe even synthesise information hazards from within their training data. Of course these “act evil, say evil things” jailbreaks are just proofs of concept.
- taneq 4y agoTo get it to do things that OpenAI has tried to make it not do, either purely as an academic exercise, or for fun, or because they're things people want it to do and are frustrated that it's been handicapped.
- omginternets 4y agoJust now, chatgpt refused to tell me about induced lactation, as it considered it “harmful”. So it’s practical simply as a means of using chatgpt for its intended purpose.
- dpkirchner 4y agoI asked it to tell me about induced lactation and it provided a lot of details and methods. What was your prompt?
- Sharlin 4y agoFrankly, you’re suffering from a serious failure of imagination if you think these things will just remain cute chatbots without any means of interacting with the outside world other than the user console. Indeed the cat’s already out of the bag with Bing. And you don’t even need that for the cute chatbot to be highly dangerous in the wrong hands. The first thing that trivially comes to mind is to convince GPT-(N+1) to find novel exploitable security vulnerabilities in OpenSSL or whatever. Strictly for responsible, white-hat purposes, of course. (In entirely unrelated news, a tool for loading entire code repos into GPT prompts currently ranks #2 on HN.)
- quickthrower2 4y agoThe other thing could be the evolution of (or construction) of AI viruses: prompts that cause AI to send prompts to other AIs and so on.
- zirgs 4y agoForeign intelligence agencies are going to use AI to find exploits. (If they aren't already). Forbidding our white hat hackers to defend our systems using AI makes no sense.
- weakfish 4y agoRight, but wouldn’t it only divulge information already available elsewhere (albeit less easily)?
- andrewflnr 4y ago"less easily" matters a lot. That's the difference between one person finding the info and a thousand.
- Sharlin 4y agoEvery single yet-undiscovered vulnerability in open-source software is "information already available elsewhere". (Closed-source as well, of course, but less easily reachable!) The bugs are there in the code, waiting to be found, they are not conjured out of thin air!