3 ms·
The actual paper is here: http://proceedings.mlr.press/v139/yeats21a/yeats21a.pdf http://proceedings.mlr.press/v139/yeats21a/yeats21a.pdf Summary seems to be t
by owlbite 5y ago
The actual paper is here: http://proceedings.mlr.press/v139/yeats21a/yeats21a.pdf http://proceedings.mlr.press/v139/yeats21a/yeats21a.pdf
Summary seems to be that ambiguity attacks can be resisted by regularization such that basin of attraction for each class is larger and it is harder to subtly nudge inference form class A to class B. The complex representation makes the regularization better.
Mechanism for this is essentially that by encoding real to complex as y = { sin(x), cos(x) } and going back again by taking the absolute value, the fallout is that a descent direction constraint manifests such that {the step in the activation descent direction}*2 + {step in regularization space}*2 = 1, so large steps are either modifying the activation or are modifying the regularization, not both, so the result is a more robust training direction.
- im3w1l 5y agoSo I tried thinking about what this really means, and I guess to some extent it measures a weighted "consensus" of inputs? First consider the simplified case of real weights. Then it will simply detect if they inputs are similar numbers or not. Similar numbers -> similar angles -> addition of vectors all pointing in same direction -> high magnitude. With complex weights there is a also a rotation of the vectors which complicates the picture (before 2.1, 1.9, 1.8 would be strong agreement, with rotation it might instead be something like 2.1, -0.1, 0.5) but I think you can still somehow thing of it as measuring if there is a consensus or not. Now since the inputs are features from a previous layer I guess what it is looking for is stuff like "Feature A should be 1.6 more strongly activated than Feature B".