9 ms·
Adversarial machine learning is an extremely hard problems space. If you have a single static model and your adversaries can react to it, then it's an almost i
by mFixman 3y ago
Adversarial machine learning is an extremely hard problems space.
If you have a single static model and your adversaries can react to it, then it's an almost impossible fight unless you are willing to give a ton of false positives and block a lot of perfectly valid prompts.
Microsoft cannot train a new model every second, but attackers can change their strategy depending on the Chatbot's answers; any "Is this user trying to access your prompt?" would be broken easily.