3 ms·
Direct and Indirect AI Injections and Their Implications
- yosito 4y agoIt seems like a good way to mitigate these attacks is to train a separate "supervisor" AI that watches all conversations for things like content policy violations and prompt injections. The supervisor AI wouldn't be a chat-based LLM, it wouldn't ever change its behavior based on prompts. Its job would basically just be to watch the chat and either approve or deny the input or output. If it did block input or output, the user could get a message in the UI explaining that a supervisor blocked the chat. For infractions too severe, it could even terminate the chat.
- kerng 4y agoVery interesting thought on how to mitigate this, because I think a solution like with parameterized queries isnt possible - at least with my current understanding (the attack is more of a "social engineering" attack on the AI). Regarding the supervisor AI, in theory it would be vulnerable to the same attack but probably more difficult to perform. One could even have multiple supervisors (with different sensitivity levels or focus areas) to get a vote on the content I guess. Interesting problem space.