3 ms·
It opens a can of worms for them if they do consider prompt injection a bug because there's ultimately no defense. If they accept this, there are instantly hund
by jdiff 3mo ago
It opens a can of worms for them if they do consider prompt injection a bug because there's ultimately no defense. If they accept this, there are instantly hundreds of other moles they now have to whack or pay out for.
Or dismiss them all as social engineering and keep it moving.
- orbital-decay 3mo ago>because there's ultimately no defense Kind of? It's not fixable as a spherical class of attacks in vacuum, but you can do a lot to mitigate particular cases, and in most cases you can patch unnecessary side channels for the injection to reach the context in an unintended way.
- keepamovin 3mo agoIsn’t it trivially fixable by having a monitor LLM? The monitor just reviews each turn pair and asks, “Is this conversation being manipulated via prompt injection?”
- zapkyeskrill 3mo agoIs it? Or does it just make it multi dimensional? As in, prompt now need to anticipate there being a monitor and instruct that one too, indirectly.
- keepamovin 3mo agoRight - but that sounds too intractable to hold up. See my other comment, I feel a chain of monitors defeats it. But hey! Who knows?
- jdiff 3mo agoAn n-deep chain of monitors doesn't really have any defense that an (n-1)-deep chain of monitors has. None of them have the capacity to separate data and instructions. All you're doing is (in some ways) giving the model more rolls of the dice to catch what's going on, but the kind of dice and the needed values to roll are in the attacker's hands as much as yours.
- keepamovin 3mo agoNah, I wanna see the data on it. Run the experiment
- jdiff 3mo agoNo, I'm satisfied with my own experience. If you would like to see data, take action towards those ends.
- keepamovin 3mo agoYes
- orbital-decay 3mo agoSuch LLM would be susceptible to injections itself, even if it's not instruction-tuned (or it would be too dumb to work as a reliable guardrail). Chain injections are trivial enough, current black box style agentic systems are easily reverse engineered in practice if you have any understanding. You can mitigate it in a way similar to the security of any human organization, but fundamentally it's a cat and mouse game, just like in any human organization.
- keepamovin 3mo agoI understand that sounds possible in theory but honestly cannot conjure an example. Care to? Even if, doesn't the monitor separation make it immune enough? I feel this is one of those "exponential" benefits things - if one is not enough, add more! A chain of monitors - "Am i being manipulated?" "Am I being manipulated?" and so on. At some point, the monitors win (and maybe approximate consciousness processes), and the prompts lose. It's interesting how close it is to "social engineering" and security/espionage organizationally. I guess the crucial difference is that incentives can be more rigorously controlled.
- RugnirViking 3mo agoHave you ever played Gandalf? https://gandalf.lakera.ai/baseline https://gandalf.lakera.ai/baseline I can assure you its very possible to win with a vast array of techniques. It doesn't prove anything, but is a fun exercise in this sort of issue.