5 ms·
This is probably not sufficient. Even if the model develops two separate pathways of data processing, eventually information has to flow beyond the "security bo
by greshake 4y ago
This is probably not sufficient. Even if the model develops two separate pathways of data processing, eventually information has to flow beyond the "security boundary". Determining whether information is hazardous down the line is going to be undecidable in the general case. Can you mitigate individual attacks? Yes, but only one working prompt can lead to a whole mess of severe issues we outline in the paper. If these LLMs are going to be your personal assistant and gatekeeper to all of your data, how much risk is acceptable?
Also, if that was the solution OpenAI would have already implemented it, right?
- sillysaurusx 4y ago> Also, if that was the solution OpenAI would have already implemented it, right? Hah. One of my most surprising discoveries in ML is that the answer to this sort of question is "Probably not!" But it took a couple years to start trusting myself and stop thinking that the pros are omniscient. In reality it's a huge undertaking to try an experimental idea like that. You have to plan for it (in the tokenizer design, in the reinforcement feedback cycle, etc) and old models can't easily be retrofitted with new tokens. This is why one of OpenAI's biggest mistakes was that they didn't reserve ~128 tokens to have special user-defined meanings, for exactly this type of scenario. Now we're stuck with their original encoder. Mitigating this specific attack is enough, the same way that mitigating SQL injection attacks is enough. You can argue "Is using an SQL database an acceptable risk?" but the answer is "Yes, as long as you sanitize your inputs." This seems to have happened because they didn't expect that users would be able to dump the original Bing prompt -- nobody was supposed to know that [system] had a special meaning. But once the model revealed that, it was a matter of time till a clever person like yourself realized that they can insert [system] into webpages.
- greshake 4y agoThis is not the same. Prepared statements eliminate SQL injections. "Maliciousness" of these inputs is well defined and can be decided by a computer. It would not be acceptable practice to "mitigate" SQL injections by blacklisting queries every time you detect a new malicious one. As these models get larger and more complex, more such opportunities for manipulation could open up, not less.
- sillysaurusx 4y ago> It would not be acceptable practice to "mitigate" SQL injections by blacklisting queries As a former pentester, this is exactly how SQL injections were mitigated in practice. Specific characters were escaped. The most surprising example was Citadel's webapp, which went from "typing ' can inject arbitrary SQL" to bulletproof within 3 days of me hammering on it. They didn't have time to switch to prepared statements, and didn't even realize SQL injection was a problem in the first place. We're at the era of "Nobody realized SQL injection was a problem." Give it time. There are solutions here. I think "escaping" the webpage by boxing it in with special tokens that can't be generated by webpages will work fine. If it doesn't, the more general solution is to have two separate context windows, one for instructions, and one for data. The model would need to be trained not to obey anything in the data window; it only informs the model of knowledge it wasn't explicitly trained on (e.g. webpages). Then you feed the website into the data window instead of the context window. Problem solved. To put it another way, which feels more likely? That 20 years from now, we'll still have zero ways of mitigating these attacks? Or that the attacks become progressively harder and harder to pull off, just like every other attack in the history of software? By the way, you should really test whether your injection still works if you remove [system] from the injection string. If you can't make bing talk like a pirate without [system], then you're SOL -- Bing's solution is to simply strip out [system] from all website data before inserting it into the context window. Kudos for using this as an opportunity to demonstrate a bunch of other types of potential vulns, though. But those other vulns need to be demonstrated in practice. Have you shown that they actually work on Bing / ChatGPT? I.e. any attacks that don't rely on [system].
- matthewdgreen 4y agoFuture generations of these attacks are going to be discovered by adversarial ML, not by humans. Models will be trained to exploit other models, in the same way that game-playing models are trained to play themselves. Unless we develop a stronger theory of what’s possible, defenses will be identical. Human beings saying things like “have you discovered any attacks” to other human beings is going to be meaningless data, and as quaint as writing large programs in assembly.
- 4y ago