4 ms·
Isn’t it trivially fixable by having a monitor LLM? The monitor just reviews each turn pair and asks, “Is this conversation being manipulated via prompt injecti
by keepamovin 3mo ago
Isn’t it trivially fixable by having a monitor LLM? The monitor just reviews each turn pair and asks, “Is this conversation being manipulated via prompt injection?”
- zapkyeskrill 3mo agoIs it? Or does it just make it multi dimensional? As in, prompt now need to anticipate there being a monitor and instruct that one too, indirectly.
- keepamovin 3mo agoRight - but that sounds too intractable to hold up. See my other comment, I feel a chain of monitors defeats it. But hey! Who knows?
- jdiff 3mo agoAn n-deep chain of monitors doesn't really have any defense that an (n-1)-deep chain of monitors has. None of them have the capacity to separate data and instructions. All you're doing is (in some ways) giving the model more rolls of the dice to catch what's going on, but the kind of dice and the needed values to roll are in the attacker's hands as much as yours.
- keepamovin 3mo agoNah, I wanna see the data on it. Run the experiment
- jdiff 3mo agoNo, I'm satisfied with my own experience. If you would like to see data, take action towards those ends.
- keepamovin 3mo agoYes
- orbital-decay 3mo agoSuch LLM would be susceptible to injections itself, even if it's not instruction-tuned (or it would be too dumb to work as a reliable guardrail). Chain injections are trivial enough, current black box style agentic systems are easily reverse engineered in practice if you have any understanding. You can mitigate it in a way similar to the security of any human organization, but fundamentally it's a cat and mouse game, just like in any human organization.
- keepamovin 3mo agoI understand that sounds possible in theory but honestly cannot conjure an example. Care to? Even if, doesn't the monitor separation make it immune enough? I feel this is one of those "exponential" benefits things - if one is not enough, add more! A chain of monitors - "Am i being manipulated?" "Am I being manipulated?" and so on. At some point, the monitors win (and maybe approximate consciousness processes), and the prompts lose. It's interesting how close it is to "social engineering" and security/espionage organizationally. I guess the crucial difference is that incentives can be more rigorously controlled.
- RugnirViking 3mo agoHave you ever played Gandalf? https://gandalf.lakera.ai/baseline https://gandalf.lakera.ai/baseline I can assure you its very possible to win with a vast array of techniques. It doesn't prove anything, but is a fun exercise in this sort of issue.