4 ms·
this seems to be an inherent flaw of the current generation of LLMs as there's no real separation of user input. you can't "sanitize" content before placing it
by bstsb 1y ago
this seems to be an inherent flaw of the current generation of LLMs as there's no real separation of user input.
you can't "sanitize" content before placing it in context and from there prompt injection is almost always possible, regardless of what else is in the instructions
- hiatus 1y agoIt's like redboxing all over again.
- reaperducer 1y agoIt's like redboxing all over again. There are vanishingly few phreakers left on HN. /Still have my FŌN card and blue box for GTE Links.
- Fr0styMatt88 1y agoGreat nostalgia trip, I wasn’t there at the time so for me it’s second-hand nostalgia but eh :) https://youtu.be/ympjaibY6to https://youtu.be/ympjaibY6to
- lightedman 1y agoSomewhere in storage I still have a whistle that emits 2600Hz.
- soulofmischief 1y agoDouble LLM architecture is an increasingly common mitigation technique. But all the same rules of SQL injection still apply: For anything other than RAG, user input should not directly be used to modify or access anything that isn't clientside.
- simonw 1y agoHave you seen that implemented yet?
- Emiledel 1y agoI've shared a repo here with deterministic, policy driven routing of user inputs so as to operate with it without influencing agent decisions (though it's up to tool calls to take precautions with what they return) https://github.com/its-emile/memory-safe-agent https://github.com/its-emile/memory-safe-agent The teams at owasp are great, join us !
- soulofmischief 1y agoI'm very curious how OWASP has been handling LLMs, any good write-ups? What's the best way to get involved?
- soulofmischief 1y agoOh hey Simon! I independently landed on the same architecture in a prior startup before you published your dual LLM blog post, though unfortunately there's nothing left standing to show since that company experienced a hostile board takeover, the board squeezed me out of my CTO position in order to plant a yes man, pivoted to something I was against, and then recently shut down after failing to find product-market fit. I still am interested in the architecture, have continued to play around with it in personal projects, and some other engineers I speak to have mentioned it before, so I think the idea is spreading although I haven't knowingly seen it in a popular product.
- simonw 1y agoThat's awesome to hear! I was never sure if anyone had managed to get it working.
- soulofmischief 1y agoNot quite the same, but OpenAI is doing it in the opposite direction with their thinking models, hiding the reasoning step from the user and only providing a summarization. Maybe in the future, hosted agents have an airlock in both directions. > ... in the future we may wish to monitor the chain of thought for signs of manipulating the user. However, for this to work the model must have freedom to express its thoughts in unaltered form, so we cannot train any policy compliance or user preferences onto the chain of thought. We also do not want to make an unaligned chain of thought directly visible to users. > Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users. Source: https://openai.com/index/learning-to-reason-with-llms/ https://openai.com/index/learning-to-reason-with-llms/
- drdaeman 1y agoDo you mean LLMs trained in a way they have a special role (i.e. system/user/untrusted/assistant and not just system/user/assistant), where untrusted input is never acted upon, or something else? And if there are models that are trained to handle untrusted input differently than user-provided instructions, can someone please name them?
- soulofmischief 1y agoSimon W has a nice write-up on it. https://simonwillison.net/2023/Apr/25/dual-llm-pattern/ https://simonwillison.net/2023/Apr/25/dual-llm-pattern/
- deleted 1y ago[deleted]
- normalaccess 1y agoLLMs suffer the same problems as any Von Neumann architecture machine, It's called "key vulnerability". None of our normal control tools work on LLMs like ASLR, NX-Bits/DEP, CFI, ect.. It's like working on a foreign CPU with a completely unknown architecture and undocumented instructions. All of our current controls for LLMs are probabilistic and can't fundamentally solve the problem. What we really need is a completely separate "control language" (Harvard Architecture) to query the latent space but how to do that is beyond me. https://en.wikipedia.org/wiki/Von_Neumann_architecture https://en.wikipedia.org/wiki/Harvard_architecture AI SLOP TLDR: LLMs are “Turing-complete” interpreters of language, and when language is both the program and the data, any input has the potential to reprogram the system—just like how data in a Von Neumann system can mutate into executable code.
- fc417fc802 1y agoIsn't it more akin to SQL injection? And would a hypothetical control language not work in much the same way as parameterized queries?
- normalaccess 1y agoThe more I looked into it it's not just the control language itself we need but a way of querying the model that is completely orthogonal to human language. But I think that would be impossible because as the newer models grow they would soon understand the control language re-blurring the line between the control language and output language. "Speaking" in a new way will not be outside it's ability to pattern match. When your fundamental compute block is language itself (and not a subset) you bounce into the limits of our understanding of language and cognition. It's a new Tower of Babel we are building by pouring all of humanities records into a mold and hoping a tower to heaven pops out the other side.
- username223 1y agoThis. We spent decades dealing with SQL injection attacks, where user input would spill into code if it weren't properly escaped. The only reliable way to deal with SQLI was bind variables, which cleanly separated code from user input. What would it even mean to separate code from user input for an LLM? Does the model capable of tool use feed the uninspected user input to a sandboxed model, then treat its output as an opaque string? If we can't even reliably mix untrusted input with code in a language with a formal grammar, I'm not optimistic about our ability to do so in a "vibes language." Try writing an llmescape() function.
- LegionMammal978 1y ago> Does the model capable of tool use feed the uninspected user input to a sandboxed model, then treat its output as an opaque string? That was one of my early thoughts for "How could LLM tools ever be made trustworthy for arbitrary data?" The LLM would just come up with a chain of tools to use (so you can inspect what it's doing), and another mechanism would be responsible for actually applying them to the input to yield the output. Of course, most people really want the LLM to inspect the input data to figure out what to do with it, which opens up the possibility for malicious inputs. Having a second LLM instance solely coming up with the strategy could help, but only as far as the human user bothers to check for malicious programs.
- whattheheckheck 1y agoSame problem with humans and homoiconic code such as human language
- whatevertrevor 1y agoIn your chain of tools are any of the tools themselves LLMs? Because that's the same problem except now you need to hijack the "parent" LLM to forward some malicious instructions down. And even if not, as long as there's any _execution_ or _write_ happening, the input could still modify the chain of tools being used. So you'd need _heavy_ restrictions on what the chains can actually do. How that intersects with operations LLMs are supposed to streamline, I don't know, my gut feeling is not very deeply.