4 ms·
FYI, Gabor-like filters pop out from doing ICA (i.e. Independent Components Analysis), not PCA. While PCA looks for orthogonal vectors onto which the data's pro
by psb217 14y ago
FYI, Gabor-like filters pop out from doing ICA (i.e. Independent Components Analysis), not PCA. While PCA looks for orthogonal vectors onto which the data's projection is normally-distributed (among other properties), ICA, roughly speaking, looks for a set of orthogonal vectors onto which the data's projection has maximal kurtosis (among other properties).
It is the kurtosis-maximization of ICA that tends to produce filters mimicking those found in (early layers of) visual cortex. Hence, the production of such filters by techniques like "sparse coding" and "sparse autoencoders", which explicitly pursue highly-kurtotic representations of the training data. PCA, on the other hand, tends to produce checkerboard (i.e. 2d sinusoidal) filters of various frequencies when trained on "natural image patches".
See: "The 'independent components' of natural scenes are edge filters" by Bell and Sejnowski, 1997.
- mturmon 14y agoThanks for the reminder. I was thinking of this 1991 paper, which I ran into a long time ago: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.41.192 http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.41.1... They used a (linear) "neural network" with gradient descent training that implemented PCA (kind of an iterative graham-schmidt process), and got Gabor-like filters. I think a lot of people have done similar experiments, with varying results.
- psb217 14y agoI hadn't seen that paper before; thanks for the reference. I read through it and saw that they were reweighting the sampled image patches with a Gaussian mask prior to learning, which explains how they got Gabor-like filters. The masking effectively forced the learned filters to have localized receptive fields, while locality/nonlocality is generally one of the (visually) clearer differences between filters learned with ICA/PCA. In other words, the Gaussian-modulated part of Gaussian-modulated sinusoids was built into their learning process, rather than appearing as an emergent property. I also chuckled a bit when they described how computing eigenvectors for 4096x4096 matrices was "beyond reasonable computation".