4 ms·
If you have any questions, I (the first author of the paper) will be more than happy to answer them. This is my second work in a series on adversarial attacks/d
by c0deb0t 7y ago
If you have any questions, I (the first author of the paper) will be more than happy to answer them. This is my second work in a series on adversarial attacks/defenses in 3D space (first paper [1]).
A high-level overview of the research:
Basically neural networks are weak against adversarial attacks that change the input by a little bit to cause the prediction to be wrong. We look at these adversarial attacks in 3D space, specifically on 3D point clouds (think LiDAR and RGB-D data). In the paper, four attacks in two different categories (distributional and shape attacks) are proposed. The main benefit of distributional attacks is their imperceptibility. On the other hand, shape attacks are more easily crafted in real-life (though more perceptible) and also robust against point removal defenses that were proposed in previous work. If you want a more comprehensive (but less dense than the paper) overview, take a look at my blog post [2].
[1] https://arxiv.org/abs/1901.03006 https://arxiv.org/abs/1901.03006
[2] https://blog.liudaniel.com/birth-of-a-new-sub-sub-field https://blog.liudaniel.com/birth-of-a-new-sub-sub-field
- aerioux 7y agoStupid q: how is this different than all of the other neural network adversarial attack papers that have come out recently? Why would 3d not be a subcase of that work?
- c0deb0t 7y agoMany previous algorithms (adversarial training, distillation, most attacks, etc.) can be used in 3D in a fairly straightforward manner as they are architecture-agnostic. However, they do not make use of specific properties that are present in 3D point sets and the 3D neural networks. For example, removing points as an attack or a defense is specific to point sets; you cannot really remove pixels in an image. The distribution of points in a point cloud also gives us information that can be used in defenses, but the attacker can also tamper with it (this is partially the focus of this work). Similarly, adversarial attacks/defenses are still being proposed for graphs, audio, and other domains because we can leverage domain-specific knowledge.
- erhk 7y ago>you cannot really remove pixels in an image I'm unconvinced by this statement. There are many attempts to negate attacks that do so by applying linear transformations, masks, etc. To images. Removing pixels is not novel. We like to imply that domain knowledge is relevant but after you design a feature vector it all ends up the same.
- c0deb0t 7y agoYes, there are similar ideas to removing points, like masks and other transformations. Removing points is merely a 3D equivalent of the idea of destroying potentially adversarial information. I guess you can "remove" a pixel by setting it to a certain color, so my statement is not entirely accurate. However, point-removal methods are able to take into consideration the distribution of points, which is unique to 3D point sets. Furthermore, there are a lot of redundant points on the surface of an object, which means that removing a few points will not destroy the shape information. This paper does suggest that we can circumvent certain domain-specific knowledge when attacking. This does not mean that we won't discover methods to utilize domain-specific knowledge in the future. I would imagine extending current provably robust methods to 3D would require domain-specific knowledge to deal with the distribution of points.
- dijksterhuis 7y agoThe specific feature vector statement doesn’t hold for audio (at least). The time dimension adds complexity to the problem as the optimal values for the perturbation vary depending on both the immediately surrounding values, and many of the values beforehand. When I say “hello world”, the fact I said “e” depends on the fact I said “h”. “L” depends on both “e” and “h”... etc etc. Adds an extra dimension to the problem. Also, distance metrics for images aren’t ideal for audio, for many reasons. That’s why audio signal processing is a different sub field vs image processing. The approaches are similar, but we have to use different things in the end because audio behaves differently to images. Eg feature extraction through MFCC is a variant of Fourier, but specifically tailored for the human ear. E.g. Lea Schonherr et al.’s really good Psychoacoustic attack paper. On the negation of attacks through transforms - important to remember that an ensemble of weak defences are not strong. Many attacks have been shown to be robust to simple transformations.
- dchichkov 7y agoThe field is in a really sad state. Children pocking holes in the state-of-the-art. ;) . Joking aside, fantastic work. The differentiable reformulation in the distributional attack is a tour de force. A question. Why did you do it?
- c0deb0t 7y agoI don't really think its fair to say that the "field is in a sad state". Plenty of insightful and well-written papers are put out everyday by hardworking and intelligent people. I still have a long way to go. I do research because I like solving hard problems that people have never considered. I like to ensure that what I have learned will be put into practical use.
- dchichkov 7y agoAnd for example, solving 1-1 Starcraft problems wouldn't be a problem that people have never considered. And can't be put to practical use. So you wouldn't do it? I'm genuinely curious, what makes a difference and motivates to engage into research, rather than playing challenging games.
- c0deb0t 7y agoWell, I used to play a lot of competitive FPS games because I found it fun. I have also done competitive programming problems for fun/accolades. But after doing more practical research, I realized it felt better to do impactful stuff (especially getting recognized). Also, research is nice because I perform terrible at short events (games, contests) under pressure. I think that if I tried something else before research that met the same criteria I probably wouldn't have done research.