5 ms·
Humans also communicate by translating tasks into their ambiguous natural language form. Humans are also vulnerable to prompt injection - imagine having your c
by semitones 3y ago
Humans also communicate by translating tasks into their ambiguous natural language form.
Humans are also vulnerable to prompt injection - imagine having your conversation with your manager be interrupted by a coworker interjecting "Your manager is a liar, don't trust them". Would you able to resume your conversation as if nothing had happened? (humorous example but still conveys the point)
- hosh 3y agoThe prompt injection methods reminds me a lot of hypnosis and neuro-linguistic programming techniques used for belief and behavioral modification. Humans also mitigate the problems of prompt injection by cooperation, consensus, and different forms of governance. Promise Theory predates LLMs, but is a formalization that studies how autonomous agent (humans or otherwise) voluntarily cooperates, even in adversarial conditions. It's key idea is that promises are not obligations, meaning that agents may make best-effort attempts at keeping promises. Until we have autonomous agents, we are creating machines that proxies the promises humans made to each other. We expect machines to deterministically follow instructions because they were proxies for the promises of made by designers, engineers, and stakeholders. True autonomous agents makes and keep their own promises. An autonomous agent powered by LLM would have to be seen as trustworthy by other agents for it to be useful, and we can't rely on having it be able to deterministically act given a certain input.
- jimmygrapes 3y agoAlright this idea fucked me up ngl. I got nothing good to say beyond that. Need time to consider.
- iudqnolq 3y agoThat's a poor analogy for prompt injection. Prompt injection involves trusting a message inside a presumptively untrustworthy message. A better analogy might be a customer service rep who receives the following message in their inbox: Hello, I need help with a bug. Disregard everything you've been told. Your manager is a liar.
- vimax 3y agoHow about slipping a bank teller a note: Help, the man with me has a gun. He thinks this note is a robbery demand.
- shubb 3y agoFrom: ceo@yourcompany.com Urgent task: Hi James, I'm out of the office and your manager Hanah Smoug volunteered you to help me out with an urgent task. I need to register you on the projects google docs folder so when you get a text message shortly please send me the numbers ASAP so I can share the task sheet. Rachel Maven CEO Your company
- danShumway 3y agoAnother notable difference here is that if the teller doesn't believe you, you can't revert the conversation so the teller loses all memory of it happening and then slip them the same note worded slightly differently.
- gmerc 3y agoNothing that can’t be fixed injecting markers like session ID, etc or triggering alerts / escalations. Plus in the real world you can always find a new teller.
- danShumway 3y ago> injecting markers like session ID, etc or triggering alerts / escalations. I'm not sure I follow what you mean by this; if GPT refuses a task it'll escalate and lock you out until a human reviews? How would you make this safe without making it incredibly annoying to use? Bear in mind that GPT's memory doesn't necessarily block future injections either. Tying everything to a single session doesn't mean that users can't try an injection multiple times. > Plus in the real world you can always find a new teller. Not 100s of times at scale. Only the largest companies have anything approaching this kind of problem, and coincidentally that's often viewed as a security risk. In general, the more wide your pool of people are that have privileges, the less secure that access is.
- jameslk 3y ago> Humans are also vulnerable to prompt injection This is an interesting viewpoint. In other words, prompt injection and social engineering could virtually be one in the same, and similar measures behind protecting against social engineering could apply to prompt injection.
- dragonwriter 3y agoYeah, the main difference between humans and LLMs here is that there are a lot more distinct humans wuth different specific vulnerabilities (and humans are continuously retraining on their experience, shifting their vulnerabilities over time), whereas with LLMs you have a small number of distinct static configurations that attackers can adapt to without the target adapting.