3 ms·
I'm not sure I understand how image grammars (whatever that is exactly) suddenly pop up as a solution after such a long introduction. I could not find any evide
by teh 11y ago
I'm not sure I understand how image grammars (whatever that is exactly) suddenly pop up as a solution after such a long introduction. I could not find any evidence or literature about them in relation to learning.
The author is stating a hypothesis but no way to test it. I'm not sure what point he's trying to make.
- deleted 11y ago[deleted]
- bra-ket 11y agowell if you take "sparse autoencoders"(http://web.stanford.edu/class/cs294a/sparseAutoencoder.pdf http://web.stanford.edu/class/cs294a/sparseAutoencoder.pdf) they basically give you principal components of the data, so from blobs of raw pixels you get a "dictionary" of lines in different orientation and size, like an edge detector. This is also similar to how the famous Fourier transform works ( http://en.wikipedia.org/wiki/Fourier_transform http://en.wikipedia.org/wiki/Fourier_transform), it decomposes the raw signal like speech into its base forms or "harmonics". And if you add another layer (to a stacked autoencoder) it will extract higher level forms(e.g. basic shapes like triangles, ellipses etc) and so on until you get a dictionary of different forms, with each layer kind of compressing the signal into more compact summary. At the higher layer you can arrive at the abstract "chair" or "cat" representation which is based on all these lower level forms: shapes, lines and dots. Then once you got this dictionary of image "words", next thing is to infer how these words interact with each other, i.e. build a grammar (it's also called "grammar induction" in natural language processing http://techtalks.tv/talks/deep-learning-of-recursive-structure-grammar-induction/58089/ http://techtalks.tv/talks/deep-learning-of-recursive-structu...). By learning a grammar you essentially define a concept of "chair" or "cat" at a higher level of abstraction by determining how these forms relate to their world (i.e. to other forms), e.g. you can determine that "cat sits on a chair" is a legal phrase in the grammar and "chair sits on a cat" is not. So extracting a grammar (visual or linguistic) from training data is equivalent to restricting the system to common sense reasoning which operates on concepts in terms of "production rules" of the grammar: http://en.wikipedia.org/wiki/Production_(computer_science) http://en.wikipedia.org/wiki/Production_(computer_science), and it is a basic goal of AI.