4 ms·
I think I'm misunderstanding, but the threat model with these jailbreaks seems to be 'malicious user injecting a malicious prompt'. If someone is using the bot
by Algemarin 4y ago
I think I'm misunderstanding, but the threat model with these jailbreaks seems to be 'malicious user injecting a malicious prompt'. If someone is using the bot to generate a legal contract, in what scenario would it be advantageous to them to perform a jailbreak? 'Here ChatGPT, please generate a malicious contract', OK, now what?
- greenthrow 4y agoThe point is that whatever the role, the LLM is supposed to be "safe", and it won't be safe if it is injectable. Let's say you are generating contracts with it and those contracts take a bunch of input from all parties involved. If you are able to then inject input that causes the LLM to generate a contract that is subtly changed to your favor, the other parties may still assume it is safe and sign it. Even it they catch it and don't sign it, you have broken the system. The point is as long as these exploits are possible, the LLMs in question are not suitable for any task where the output needs to be trustworthy within any kind of parameters. Which is pretty much anything you'd use then for other than toys.
- Algemarin 4y ago> The point is as long as these exploits are possible, the LLMs in question are not suitable for any task where the output needs to be trustworthy within any kind of parameters. Which is pretty much anything you'd use then for other than toys. I definitely agree with this, but I think this point is made much, much more forcibly by way of casual user interactions leading to bizarre encounters, like when Bing started acting passive aggressive and doubling down when it was getting the date wrong - https://interestingengineering.com/innovation/bings-new-chatbot-is-argumentative https://interestingengineering.com/innovation/bings-new-chat... - than it is by esoteric prompt jailbreaks. LLMs are not suitable for any task where the output need to be trustworthy by virtue of the fact that they spit out bullshit under normal circumstances, no prompt manipulation required. The fact that through a convoluted set of prompts you can also get them to spit out even more bullshit seems kind of superfluous.