3 ms·
Very interesting results, but to imply that this is some kind of security vulnerability is humorous to me. Trying to prevent “prompt injection attacks” is the e
by sippeangelo 3y ago
Very interesting results, but to imply that this is some kind of security vulnerability is humorous to me. Trying to prevent “prompt injection attacks” is the equivalent of trying to string replace quotes out of your SQL query. mysql_real_escape_string() anyone?
What are people doing with these LLMs where the worst thing that can happen isn’t just the legal team saying “We don’t want our AI to be tricked into saying something controversial and a user screenshotting it”?
- Szpadel 3y agothe thing is that we don't yet know how to implement llm_escape_string()
- pixl97 3y agoDidn't MGM just have a human_agent_escape_string() attack?
- joshuanapoli 3y agoThe instructions and data, input and output are all mixed into a single domain. So we end up with similar weaknesses as simple processors. And similar to how we all get tripped up by phishing. So we need memory protection. Or MFA.
- kromem 3y agoWe broadly do. It's called multiple passes. It's just 2-3x more expensive to run a prompt and then the prompt and response though a secondary fine tuned model to detect injection when the cost is so low impact. So no one is doing it.
- mechagodzilla 3y agoAn obvious concern seems like any LLM-style 'agent' that can use your credentials to act on your behalf (financial, email, file access, etc.)