3 ms·
> Do you think that the idea of AI poisoning generalizes, though (even if it would require very different algorithms) Yes, these are all different classes of a
by declaredapple 3y ago
> Do you think that the idea of AI poisoning generalizes, though (even if it would require very different algorithms)
Yes, these are all different classes of adversial attacks and I don't think any model could resist everything, assuming you have direct access to the weights (although that's generally not a hard requirement)
> or do you think that some models may be "safe"?
I don't think any single set of model weights can be "safe". I suspect adding more adversaial data to the training, particularly of the text encoders might make them more resistent.
I unfortunately don't have high hopes that artists have a great chance of defending their work with these methods, nearly all of the existing ones are easily evadable, and I suspect more adversarial resistant methods will be implemented.
On top this, I think resisting adversial inputs will be a focus for many, not because they want to maliciously copy data, but because of the implications of multimodal models being attacked. Imagine a billboard that will be recognized as instructions to do XYZ (has been demoed with GPT4V)