4 ms·
I don't mean to be super negative, but because of the general tone early in the article and some sloppy notation, I never finished reading. I think the goal of
by TravisDick 11y ago
I don't mean to be super negative, but because of the general tone early in the article and some sloppy notation, I never finished reading. I think the goal of an article like this should be to give a high-level intuitive explanation for some technical result, rather than sounding smart or complicated.
First, it is a little weird to me to talk about "old-school ML" as learning maps from inputs to hidden features. That seems neither old, nor very representative of the field of Machine Learning as a whole. It's also weird to say that RBMs and other deep learning algorithms are formulated using classical statistical mechanics. Moreover, implying that this scary-sounding formulation is the reason they are interesting seems like an attempt at sounding smart. Typically there are many ways to motivate and derive different algorithms, and it is /useful/ to acknowledge the multiple viewpoints because they often give different insights.
Second, the section about flow maps and fixed points seems to make a mess out of the notation by either being unclear or disagreeing with standard notation. What is meant by the notation "f(X) -> X"? Presumably this means something like f is a function that maps elements of the set X to elements in the set X. More standard notation for this would be something like "f: X -> X". Perhaps it means that the image of the set X under the function f is again the set X. But does that require that f be a surjective function? Confusingly, it also looks like the function f might be required to be the identity function, but given the context this is clearly not the intended interpretation.
When defining the fixed point, it seems that it would be more natural to say that x is a fixed point of f if f(x) = x. That is, x is fixed or unmoved by the function f. It turns out that for contractions (and some other functions, too), that the sequence f(x), f(f(x)), f(f(f(x))), and so on is guaranteed to converge to a unique fixed point of f. The notation f^n typically refers to the function f being applied n times, which is not the usage in the article. In the article, f^1, f^2, and so on are all identical copies of the function f. Using the standard notation, the definition of f_infty would be f_infty(x) = lim_{n -> infty} f^n(x). And, in the case of a contraction, the Banach fixed point theorem gives that f_infty is well-defined, and there exists a unique x_fix in X so that f_infty(x) = x_fix for all x in X (i.e., iterating f repeatedly converges to a unique fixed point x_fix of the function f).
These things do not necessarily mean that the article is uninteresting or uninformative or even technically incorrect. But if the author didn't take the time to make the simple things clear, then I'm not sure that I want to read the rest.
Sorry for the rant.
- sieisteinmodel 11y agoWell, there is more. E.g. abbreviating deep belief nets with DBM, which is the commonly used acronym for deep boltzmann machines. These are similar, but very different. Calling an RBM an encoder is somehow not far fetched, but there are many differences between auto encoders and RBMs. He eventually claims an RBM minimises reconstruction error, which is just plain wrong and shows that this guy has absolutely no clue what he is writing about.
- charleshmartin 11y ago'Technically' this is correct--the RBM CD algo is not minimizing this function; that's not the point. It is known that when training an RBM, the reconstruction error decreases but not monotonically; in fact it fluctuates. In the words of Hinton, 'trust it but don't use it'. http://www.cs.toronto.edu/~hinton/absps/guideTR.pdf http://www.cs.toronto.edu/~hinton/absps/guideTR.pdf (which is cited in the post as well) So in a global sense, yes, I would say that the RBM does eventually minimize the reconstruction error even though it fluctuates. I can even offer a conjecture here on why the error fluctuates ; in a discrete RG flow map, there could be finite size effects that would give log-periodic fluctuations. This is a stretch--but it is something that could be tested. I explain this idea here http://charlesmartin14.wordpress.com/2015/01/16/the-bitcoin-crash-and-how-nature-works/ http://charlesmartin14.wordpress.com/2015/01/16/the-bitcoin-... As to stacking the RBMs to form a DBN--yeah that's the point. "Hinton showed that RBMs can be stacked and trained in a greedy manner to form so-called Deep Belief Networks (DBN)" http://deeplearning.net/tutorial/DBN.html http://deeplearning.net/tutorial/DBN.html
- charleshmartin 11y agoThanks for the comments. It is helpful to have others read the blog and make suggestions and ask for clarifications. The motivations here are to (1) summarize and clarify the key ideas of the physics paper that observed this connection, and (2) set the stage for my next blog, where I try to connect Deep Learning to my idea around Spin Funnels. I will review the comments and think how to update the blog to make it more clear.