7 ms·
I quite surprised at the comments on HN so far as nobody seems to see the significance of this. Yes, the image is ambiguous. The point is that Google Cloud Vi
by ratel 8y ago
I quite surprised at the comments on HN so far as nobody seems to see the significance of this.
Yes, the image is ambiguous.
The point is that Google Cloud Vision gives an unambiguous answer of that image based on the rotation.
Transformations of an image are regularly used to improve the results of image recognition. That process fails quit dramatically if in the course of a transformation the answer given is presented with higher confidence than should be.
- tziki 8y agoEh, you use rotations if you're specifically looking for rotation invariance. Also, it's hard to tell how confused the network actually is since it's not really predicting 'probability' of a class (even though the term is commonly used). Quite often neural nets just output the most 'probable' class with an oversized probability estimate due to how the most common classification layer works.
- gambler 8y agoI'm glad that at least someone here sees the problem, but I am not surprised by the typical reaction of AI apologists in this thread. You always get at least one of the two responses: "OMG, this is amazing, it's just like humans. We're probably close to AGI." "Ha-ha, humans are stupid, so the algorithm giving unexpected result is just a proof that it's better and less biased." Here, we have both in response to the same demo. Still, I honestly don't know why some people are so biased in favor of neural nets and have zero interest in edge cases and flaws (the most interesting parts if you want to gain deeper understanding of how the algorithm actually operates). Wishful thinking, I guess.
- cubano 8y agoProbably, and I do hate being the cynic in the room, it's due to the huge salaries and interesting, specific use-case work that NN-based AI is generating these days.
- immernerstheime 8y agoyou are right probably , due to bullshit hiding knowledge from researchers , just for fun , like if there were Gods like greek gods or so , doing it for fun... its not a monopoly of cartelized evil researchers or capistalist owners factorylords, gatekeepers intentionally egoistic possesors of the science knowledge, No , that is not the case Its like it happens naturally, by forces of the natural human stupidity anyway all programming is evil so... dont know even what the fck am I doing commenting in this shitty topic but I have anyway the fckng freewill to fkcng comment... frewill always EveryOne thanks unimportant note the slogan Free Will Always EveryOne is open as in an open beer can and free as in fucking freewill and free beer at the Trio de Carnaval in Brazil carnival and common fucking domain of Nature Itself, myfriend, mor open then wikicommons, its The Fckng Free Domain Open and Public and Common you have the free will to use it freewill always EveryOne thanks But fucking remember: that shit is Not fucking my slogan , I dont have fucking slogans, dont make fucking analogies with fucking me, Dont make analogies with me!!! Fuck! not my fucking slogan! no fucking analogies with me! no analogies! fuck there is no Marco_s fucking Mark left No fucking Mark; no fucking marks; No mark! Cheers! thanks
- hn_throwaway_99 8y agoCan someone explain why this is a problem? I'm not an "AI apologist", but I would consider it a good thing that the model pegs it as a rabbit when it is in more of a "rabbit orientation" and a duck when it is in more of a "duck orientation".
- SlowRobotAhead 8y agoIn this case, a rabbit and duck are approximately the same size and danger level. So few cases where there is harm possible. What if it was AI looking at bacteria? Or scanning a roadside for IEDs? Or when a guy on a bike when turned and rotated the correct way appears to be a crosswalk paint mark of a guy on a bike? If our current AI is making different “DEFINITE” determinations based only on image rotation - there is a problem.
- mjburgess 8y agoNot sure why you're being down-voted, this is exactly the issue. The image is both a rabbit and a duck regardless of orientation, capturing the object it depicts as a single class with a confidence measure is a mistake. The magnitude of this mistake becomes apparent when you connect it to real-world decision making, and it becomes highly unsafe. As is most of ML because it only uses statistical (rather than causal) modelling of the world -- so really, it is only offering us generalised statistical associations. It cannot cope with statistical discontinuities.
- gambler 8y agoIn the real world, you don't want AI to instantly flip from 90% confidence in one direction to 90% confidence in the other direction, because it would cause erratic behavior. What would be preferable is a large zone where it gives both labels .45 score. Then you can apply higher-level reasoning based on the possibility that the object could be either of those two labels (i.e. act on the possibility of the most dangerous or most beneficial scenario of the two).
- armamut 8y ago
- rjf72 8y agoAI has become exactly like most complex issues with multiple distinct 'sides'. For whatever reason everybody is expected to have an opinion even though, of all people, it's generally safe to say < 1% have the knowledge and experiential basis to form an educated opinion. So the other 99% are mostly just picking sides semi-arbitrarily. And it's this arbitrariness that leads to overly, and inappropriately, simplistic responses to issues. This simplicity in turn also tends to drive responses that are either radically 'for' or radically 'against' something. Understanding drives nuance, and nuance drives uncertainty. The trouble with the world is that the stupid are cocksure and the intelligent are full of doubt. - Bertrand Russell Though in this case, stupid/intelligent are probably overly harsh. Intelligent people are certainly not immune to this 'trap'. In some ways they can be even more susceptible since they may themselves know very little, but that very little is still enough to put them ahead of 80% of the rest which can yield unjustified confidence. So let's just say uninformed/informed.
- Shorel 8y agoHumans reading a text in Latin letters will find that d, b, p, and q are different letters. In most fonts, the only difference is rotation and/or mirroring. Yet they represent different sounds, they have different meanings. Rotation is another data point, not something entirely independent from the data.
- dusted 8y agoI'm just saying it's as a rabbit in one orientation, and valid as a duck in another orientation. Many things belong in distinct groups depending on their orientation or other physical attributes. A bucket is, amongst other things, a bucket when standing flat on the ground with the hole facing upwards. Turn it around, and it becomes the cover for a mole trap, place it on someones head, it becomes a rain cover. Mount a light bulb inside it and it becomes a lampshade. It's still a bucket, but it's not _primarily_ a bucket in all cases, and shouldn't necessarily be classified as a bucket in all cases. Rotate the plus symbol 45 degrees and you got the letter x. And it is definitely, 100% certainly an x in one orientation and a + in another orientation.
- SilasX 8y agoThey have scale-invariant feature transforms (SIFTs[1]). I wonder if they could do rotation-invariant ones that wouldn't have a different answer depending on rotation? [1] https://en.wikipedia.org/wiki/Scale-invariant_feature_transform https://en.wikipedia.org/wiki/Scale-invariant_feature_transf...
- sjburt 8y agoSIFTs are rotation invariant. However, they are not good for classification. Anyway, the early layers of an NN should be performing an encoding that creates scale and rotation invariance though, so that later layers can classify. That's what makes this result interesting. Well that and the ambiguity matches the human ambiguity.
- jefft255 8y agoLookup spherical CNNs and tensor field networks, both are rotation invariant using spherical harmonics. They are not really used in practice however.
- randyrand 8y agoproblem? this is a feature imo. it uses the rotation I give it to help deduce ambiguous content, which is helpful.
- ratel 8y agoIt is only a feature, because you have the information about which rotation you provide and relate it to the result. If I just try to classify an image the information about which rotation the image is under might be not there at all. Just an example. Say there is a drone: The Elmer-2 that takes an image of a sign. Now depending on which angle this picture is taking from it might think it is duck season instead of rabbit season and never look twice.
- Klathmon 8y agoBut the human brain fails at this if it's rotated also... If you showed me the rabbit rotation of the picture, i'd tell you with pretty high confidence that it's a rabbit. If you showed me the duck rotation, i'd tell you with pretty high confidence that it's a duck. That's the point of this, it's an illusion. And it did give a bit of an "I don't know" answer for many of the rotations in the middle of the gif/video, which is exactly as I would expect it to, and when I pause the video at those points and glance at it, it doesn't look like much of anything to me either.
- 2muchcoffeeman 8y agoBut a human brain should only falls for it once. You have your initial reaction, realise it might be the other animal and then from there you know it’s an illusion and the rotation of the picture no longer matters.
- deleted 8y ago[deleted]
- angry_octet 8y agoIf they chose to allow a classification option of (c) optical illusion, then the system could easily learn that. In fact, you could probably make a GAN ambiguous image generator.
- 2muchcoffeeman 8y agoAn ambiguous image is not always ambiguous. Eg clever animal camouflage. A caterpillar that looks like a snake should not be classified as ambiguous. It should always be caterpillar.
- angry_octet 8y agoAh, but then you open the door to more complex and meaningful classification. For example "insect camouflaged as a stick" or the classic "couch with leopard print pattern".