4 ms·
I think the intuition behind max-pooling is that it says "something is in this local region", without overly specifically saying where it is. Intuitively, a hum
by zackchase 12y ago
I think the intuition behind max-pooling is that it says "something is in this local region", without overly specifically saying where it is. Intuitively, a human may detects an edge or intense bright light, but not care so much about precisely where it is in the field of vision. After many successive layers of max-pooling, however (if the pools are not over-lapping), even somewhat course information about locality is lost.
I believe Hinton objects to this gross loss of spatial information for two reasons:
1) Humans don't lose so much spatial information, and Hinton would like his models to ultimately capture a neurologically plausible computation.
2) It may not be necessary for object detection (Imagenet), but it would likely be important for more sophisticated tasks.
- bainsfather 12y agoHe also 'did not like' Support Vector Machines, back when they were the best method for image recognition. His reason was that SVMs were a 'dead end' - they were not a step on the path to human-level image recognition. His argument now seems pretty valid. I think he is saying the same thing about Max Pooling. Just my guess.