4 ms·
Actually I think the impossibility of using natural language instructions to LLMs to prevent prompt injection demonstrates (or will demonstrate) that no true un
by lukasb 3y ago
Actually I think the impossibility of using natural language instructions to LLMs to prevent prompt injection demonstrates (or will demonstrate) that no true understanding is happening.
- l33tman 3y agoThis doesn't really follow.. Any human can be fooled (and are fooled) by data for example. Humans fall for advertisement and fake news despite "understanding" what's going on, humans fall for cons all the time despite REPEATEDLY reading in the news exactly how the scams work and what to expect and how to guard against them. If you separate the prompt into two parts (like OpenAI does in their GPT API), with one "System" input and one "User" input, it only pushes the issues one step away. The User data input could certainly "spill over" into the System context and understanding as at some level, the System context is supposed to act or output stuff based on the User data. One fix is probably about the same as for humans - you need to almost autistically and in immutable OCD fashion learn to consider, during all actions you take, if this action seems to be bad somehow - perhaps with a monitoring AI "sub-process" if you like. I'm sure that can be manipulated as well though, so I predict layers of these will eventually be added..