4 ms·
Not clear if they had asked for "butterfly" instead of "cat" if 50/50 would have been the result. Similarly, if random perturbations influence choice, the basel
by RandomLensman 3y ago
Not clear if they had asked for "butterfly" instead of "cat" if 50/50 would have been the result. Similarly, if random perturbations influence choice, the baseline should include the noise from that.
- riwsky 3y agoThe null hypothesis is 50/50 for butterfly, as well. From our conscious perception, it’s two copies of the same picture of a vase.
- RandomLensman 3y agoThe hypothesis is, but I think that is something to experimentally verify in the case of no butterfly noise in the pictures.
- hervature 3y agoI'm not really sure what you are commenting on. They tested this across multiple pairs of categories: (sheep, chair), (dog, bottle), (cat, truck), and (elephant, clock). This isn't a phenomena related to cats. The whole point of the study is to measure the impact of the noise. The "baseline" or control here would be to to not add noise to either of the two images and arbitrarily label one "cat" and the other "truck" and see how the humans perform. It is obvious that humans cannot do better than 50/50 and any deviation is purely chance. In the perfect world, you would do this control to ensure your experimental setup is not flawed in some other way but if the experiment was done as double blind then this control study would be pretty silly.
- RandomLensman 3y ago1) It is not adding pure noise. 2) If humans when prompted tend to always see something more in one picture than the other when random noise is added, the baseline might not be 50/50 as no matter what you ask you get a systematic preference. Double blinding would not remove this.
- hervature 3y agoOk, I understand what you are saying. Essentially, you would have liked different type of perturbations tested to get a baseline effect of how much a random perturbation can get people to agree. They did do this in what they call Experiment 3. Their control perturbation is simply the adversarial perturbation but flipped left/right. They claim this was to preserve perturbation statistics. Still, I think baselining against a control perturbation is not the point of the study. My takeaway of the study is that they give a constructive way of influencing human perception of images. I understand that the concern could be something like "random perturbations of cat images make every image simultaneously less cat-like and hence more like anything else". My opinion is that Experiment 4 (making an image more cat-like or truck-like) covers concerns of this nature. Even if there are two random perturbations where one makes an image cat-like and the other more truck-like, it is completely arbitrary whether you label the perturbation as cat-like or truck-like (since they are randomly generated). That means, even if you measure a difference (and even if two random perturbations have larger differences than this construction!), you cannot control the direction. This method gives you a way to control it. Personally, I don't think this study is about measuring the influence over some baseline. It's about showing that you can indeed choose the direction of influence.
- RandomLensman 3y agoI agree that in some ways they control for certain aspects, but with human behavior I am leaning towards being old school in wanting to see baselines when there is nothing "structured" to have a better idea of baselines (and potentially also learn something about human perception). In the current study I am not fully convinced all confounders are controlled enough. (There is also the issue of prompting, but asking if it were not, e.g., a vase what else is there is experimentally too broad, I fear) A broader way could be to add random noise then ask to pick the more X-like imagine and see how that correlates with the classifier probability for X.