10 ms·
The author writes "Human children don’t need such explicit and extensive training to learn to recognize a household pet." This claim seems dubious. Study's hav
by jonbronson 9y ago
The author writes "Human children don’t need such explicit and extensive training to learn to recognize a household pet."
This claim seems dubious. Study's have shown humans can react to visual stimuli in as little as 1-3ms. If a child observes a cat in the room for only 10 seconds, that's already between 3,000 to 10,000 samples from various perspectives. While our human experience may describe this as a single viewing 'instance', our neurons are actually getting an extensive, continuous training. Is this accounted for in the literature?
- jayd16 9y agoSo we should be using 10 second videos instead of images to train AI?
- dr_zoidberg 9y agoI'm pretty sure that was the reasoning to build the ImageNet (about a million labeled images) in the first place. But labeling images is expensive, and there are hints there's more at play with human cognition. If you see a black cat and a white cat, and someone tells you there are striped colored cats, you can imagine it. And if you were to come across it, you'd instantly recognize it as a cat. Neurals nets can't do that. You can also see a lynx and recognize it as "some kind of cat". Again, neural nets are not there yet. Which is why there are people researching to find new, better algorithms that better mimic what we recognize as intelligence.
- yorwba 9y agoAre you sure about your examples of things neural nets can't do? I think GANs might be able to "imagine" striped cats, provided they have been trained on enough images to capture the space of black/white/striped objects. And a lynx being classified as a cat doesn't seem so outlandish. It has to be classified as something and cats are likely the closest in appearance. Of course these are just based on my intuition of what neural nets are capable of, so if you have examples of cases where these specific tasks were attempted unsuccessfully, I'm interested.
- mannykannot 9y agoI would be interested regardless of the outcome.
- dr_zoidberg 9y agoLet me remind you of sofas being classified as cats[0] and people being classified as gorillas[1]. You're overestimating the guesswork convnets are able to do, based on fragile training (which is still better guesswork than what previous models did). [0] http://rocknrollnerd.github.io/ml/2015/05/27/leopard-sofa.html http://rocknrollnerd.github.io/ml/2015/05/27/leopard-sofa.ht... [1] https://www.theverge.com/2015/7/1/8880363/google-apologizes-photos-app-tags-two-black-people-gorillas https://www.theverge.com/2015/7/1/8880363/google-apologizes-...
- yorwba 9y agoPeople being classified as gorillas was actually what I was thinking about regarding the lynx/cat example. The model might have been unsure about the kind of ape it was looking at, but clustering them together is its own kind of achievement.
- dr_zoidberg 9y agoThe thing is that the convnets are unable to learn about "macro structures" (or structures in general). A cat has ~4 legs, a tail and pointy ears. Gorillas are black, have a primate-y face and fur. The sofa is lacking the tail, head and pointy ears. People were missing the fur. Yet those things did not prevent the net from missclassffication (because those features weren't detected in the learning phase). Once again, children are able to see a cat and extract all that relevant information: four legs, head, tail, eyes & nose & ears with a particular shape, different than dogs, most cats fur (except for those alien-looking furless cats, of course).
- nkoren 9y agoIf you ask a child to draw a hand, they will almost always draw it with five fingers stuck straight out, widely separated. This is a view of a hand that one almost never actually sees; generally you'll have fingers clustered together, occluding each other, foreshortened, etc. So why do they draw it like that? It's because they're drawing a conceptual, represntational model of a hand, not a distilation of visual "hand" characteristics. That's the difference with human learning: it's based on representational model-making, which is not at all the same thing as pattern matching.
- MrQuincle 9y agohttps://github.com/junyanz/CycleGAN https://github.com/junyanz/CycleGAN
- dr_zoidberg 9y agoGANs and style transfer are not the same as being able to recognize and imagine changes on the representation you just learned. Also, look at the GAN examples: even the water in the background is being affected by the "horse to zebra" transfer. You can perfectly imagine a brown horse, standing in the beach, and then being told "now iamgine the horse is white" without than "instruction" affecting how you imagine the beach.
- MrQuincle 9y agoPerhaps explain what your preferred dataset is. + Brown horses on beaches + Some way to indicate "white" "horse". If you have a good idea about how we can train for your problem, I'm not so convinced that it cannot be solved.
- dr_zoidberg 9y agoNo, the problem I was pointing at is that you want to change a part of the image (horse into zebra), but the style transfer GAN learned to map pixels from one space into another. So it knows it has to change the colour of some things, and add stripes here and there, and also probably there too. But it isn't consistent with the stripes (you can see in some gifs how the pattern suddenly changes and adjusts), and it doesn't recognize (segment) the horse as the only relevant thing that has to be changed. But that's not bad per se about style transfer. It's an interesting technique, but if you want to convert all horses to zebras in an image, that seems to be a bit too general for current-generation GAN architectures. Maybe it can be improved upon, or a different, novel architecture is required, and not just something you can solve by throwing more data at it.
- MrQuincle 9y agoYes, the segmentation could be better. I still think you're not pointing out fundamental issues though. :-)
- nabla9 9y agoReacting to visual stimuli is reflexive and not directly related to learning. Human brain receives roughly 25-50 images worth of data per second, so less than 50 samples per second. (consciously we observe only 25 images per second). Short-term synaptic plasticity works on a timescale of 20ms to few minutes, so also roughly 50 times per second timescale. If I try to translate this to deep learning framework, it would mean max 500 training steps in 10 seconds per neuron. Learning to recognize cat using _unsupervised learning_ in just 10 seconds would be really impressive.
- EGreg 9y agoNoam Chomsky would argue it's because (nearly) all of us are born with some constrains in our brain designed for dealing with our world.
- nabla9 9y agoAt the level of object recognition we are discussing now, the "constraint" (more accurately learning bias) comes in the form of very advanced neural architecture that is able to learn in spatial environment. Not in some kind of pre-trained neuron weight collection. We have some very particular biases like fear of snakes or heights, but learning to recognize spatial objects is something very general.
- heavenlyblue 9y agoFear is neither prover or disproven to be a genetic trait.
- skummetmaelk 9y agoIt's not accounted for in the literature because brains do not work like computers processing 1 frame every x milliseconds.
- mannykannot 9y agoShow a child who has never seen an elephant a picture of one, and she will probably correctly identify the first one she sees.
- AndrewKemendo 9y agoDefine child. 2 year old? 5 year old? I have three kids and I can tell you that they wouldn't be able to do this reliably until probably age 5, and even then maybe.
- akud 9y agoMy 2 year old daughter could.
- mcv 9y agoWithout ever having seen one? Or with having seen pictures of one? It's true though that we can generalise from descriptions and recognise the real thing from those. If you describe an elephant as a big grey animal with big ears and a trunk they can use to grab stuff, then an adult (not a 2 year old I suspect) seeing one for the first time, will recognise it from that description. When we see a cartoonish drawing of one, we can still distill the defining characteristics from it and use it to create a description or recognise the real thing. We can recognise a very crude childish drawing of one by looking for these characteristics. We have a lot of additional knowledge that influences our image recognition, and having a big toolbox of general recognition of tons of different objects, we don't really need to train to recognise new objects anymore, because we will distill its identifying characteristics the first time we see it. Computers clearly don't look like that.
- Shorel 9y agoSome pattern recognition seems to be innate, for example chicks just hours old can recognize the shadow of a flying bird as either harmless (long neck short tail) or a predator (short neck long tail). So, while this pattern in particular doesn't apply to humans (it really doesn't?), many animals have ready-to-use pattern recognition when they are just hours or days old.
- amelius 9y agoChildren also understand that a cartoon-cat is a cat, even though they don't look very similar.
- Finch2193 9y agoDoes this mean you expect AI to be able to train on one video of one cat from different angles?
- kkylin 9y agoOff-topic, but I'd be really interested in seeing a reference for such a study. A single action potential is usually only 1-3 ms (see https://en.wikipedia.org/wiki/Action_potential https://en.wikipedia.org/wiki/Action_potential and references therein), and retina to LGN to V1 (the first two stops in the early visual pathway in mammals) takes seveeral tens of ms (off-hand cannot find a ref, sorry). If a human can react to a visual stim in anything less than that, it would seem to be because (i) they are anticipating the stim based on some prior cue, or (ii) there is some shortcut directly from the eye to motor neurons, and the rest of the brain isn't involved directly. In any case 1-3 ms still seems extremely fast. Also, a child observing a cat continuous for 10 seconds is getting highly correlated samples, not new independent instances. The effective number of samples (which I quite agree would be >1 if the child got to examine the cat from different perspectives, or the cat moves around, etc etc) should still be lower than what a putative sampling rate would suggest.
- Phemist 9y agoThe optic nerve is directly connected to the superior colliculus, which directly controls saccades and is sensitive to novel (bright) visual stimuli. If such a stimulus hits the the periphery just outside the fovea the time from stimulus onset to "fixation" should be minimized, without having to go through the "slow" ventral pathway/stream (V1 etc.). This time should still be more than 1-3 ms, but likely is below 100ms. No references, sorry. Edit: First said LGN projects SC, but pathway is even shorter than that. Edit2: The consensus on these fast, superior colliculus-guided "Express" saccade latency seems to be 80-120 ms. See: http://www.scholarpedia.org/article/Human_saccadic_eye_movements http://www.scholarpedia.org/article/Human_saccadic_eye_movem...
- nonbel 9y agoWhy is there so much expertise being shared without any refs on this topic? Are they hard to find online for some reason?
- Phemist 9y agoOne of the issues is this type of neuroscience is becoming too "basic", so it's what gets taught in classes as the truth and professors don't necessarily give references anymore. Also I was typing on my mobile phone, but have switched to my laptop now. The scholarpedia (basically peer-reviewed wikipedia) article on saccades should be pretty interesting.
- mcv 9y ago> If a child observes a cat in the room for only 10 seconds, that's already between 3,000 to 10,000 samples from various perspectives. Why does that count as 3000 - 10,000 samples? Why is it not a single sample? I don't think our brains sample image in that way. And that might be a fundamental difference between how humans process images and how we're expecting computers to do it.
- madamelic 9y agoIf the cat was perfectly still, not twitching and neither were you, it would be one sample. With both entities moving, you are getting constant, discrete samples of what a "cat" is.
- ben-schaaf 9y agoI think what the GP means is that there are no discrete samples and the information stream is instead one continuous sample.
- byebyetech 9y agoI believe Human brain also reuses abstract concepts it learn from other data. Such as eyes. So when a new animal is shown its very easy to find eyes on it and does not need another million images of that new animal to detect it's eyes. We need a SpaceX of Deep Learning, where a lot of learning is reused and linked in creating much larger web of knowledge about the world.
- outlace 9y agoBut if I give you a single image of a scene with an object you’ve never seen before, you’ll likely be able to instantly segment it, describe it, and relate it to things you do know.