8 ms·
Yeah, improving robustness from prompt injection which such techniques will help. One attack avenue that is surprisingly not discussed much is that the model i
by wunderwuzzi23 1y ago
Yeah, improving robustness from prompt injection which such techniques will help.
One attack avenue that is surprisingly not discussed much is that the model itself can be the attacker.
In that case prompt injection is not the root cause, but a misaligned/backdoored model that might invoke tools is.
So super risky use-cases should always require human oversight, but I'm worried we are already on a path of normalization of deviance.
It's sort of the unlikely worse case scenario, but Murphys law reminds us that such an attack/accident will happen one day.