3 ms·
Very interesting thought on how to mitigate this, because I think a solution like with parameterized queries isnt possible - at least with my current understand
by kerng 4y ago
Very interesting thought on how to mitigate this, because I think a solution like with parameterized queries isnt possible - at least with my current understanding (the attack is more of a "social engineering" attack on the AI).
Regarding the supervisor AI, in theory it would be vulnerable to the same attack but probably more difficult to perform. One could even have multiple supervisors (with different sensitivity levels or focus areas) to get a vote on the content I guess.
Interesting problem space.