4 ms·
https://archive.md/tpm8h https://archive.md/tpm8h It's cool they found an attack that works across models and apparently is based on commonality on the trainin
by version_five 3y ago
https://archive.md/tpm8h https://archive.md/tpm8h
It's cool they found an attack that works across models and apparently is based on commonality on the training corpora.
I don't see this as a real world concern - on the "getting the model to say something bad" side, who cares if a targeted adversarial attack can do that? If the only way it occurs is though entering a crafted gibberish string, I don't see it as a concern.
On making a system that contains a model malfunction, while I struggle to think of a real example, it should be a wakeup call that production models should be implemented so that there is no impact from a malfunction, whether a hallucination or an adversarial attack. Using a consequential model directly and without guardrails is irresponsible and shouldn't be done anyway.