3 ms·
Yeah, I think taint tracking was one of the early ideas here also. The problems is that the chat context typically is immediately tainted as for the AI to do s
by wunderwuzzi23 1y ago
Yeah, I think taint tracking was one of the early ideas here also.
The problems is that the chat context typically is immediately tainted as for the AI to do something useful it needs to operate on untrained data.
I wonder if maybe there could be tags mimicking data classification - to enable more fine grained decision making and human in the loop prompts.
Still a lot of unknowns and a lot more research needed.
For instance with Google Gemini I observed last year that certain sensitive tools can only be invoked in the first conversation turn / or until untrusted data is brought into the chat context. Then for the next conversation turn these sensitive tools are disabled.
I thought that was a neat idea. It can be bypassed with what I called "delayed tool invocation" and usage of a trigger action, but it becomes a lot more difficult to exploit.
- seanhunter 1y agoIt seems to me that the only robust solution has to be some sort of split-brain dual model where tainted data can only ever be input to a model which is only trained for sentence completion, not instruction-tuned. Untainted data is the only data that can be input into the instruction-tuned half of the dual model. In an architecture like this, any attempt to prompt inject would just find their injection harmlessly sentence-completed rather than turned into instructions and used to override other prompt instructions.
- wunderwuzzi23 1y agoYeah, improving robustness from prompt injection which such techniques will help. One attack avenue that is surprisingly not discussed much is that the model itself can be the attacker. In that case prompt injection is not the root cause, but a misaligned/backdoored model that might invoke tools is. So super risky use-cases should always require human oversight, but I'm worried we are already on a path of normalization of deviance. It's sort of the unlikely worse case scenario, but Murphys law reminds us that such an attack/accident will happen one day.