4 ms·
This article is unfortunately complete mathematical rubbish. The author appears throughout to show a strong lack of understanding about the mathematics behind
by tysam_and 3y ago
This article is unfortunately complete mathematical rubbish.
The author appears throughout to show a strong lack of understanding about the mathematics behind what Hinton was saying and the math behind LLMs, and tries to rebut it with casual, non-mathematical examples from their personal life from an entirely different problem domain (!!!!). They then have the gall to say about the most-cited ML researcher of all time: "Maybe Hinton’s problem in understanding this is that he’s just too logical!" No, Hinton's problem in understanding it is that he actually correctly understands the information theory behind what's happening in LLMs. He sort of founded the modern field and has been doing this for, what, four decades?
Let me explain to you what Hinton is implicitly saying here behind his words, as best as I understand it. Every language process can be interpeted as a tokens, in our case discrete. This process is generated under a system where the one driving variable is time, and it is autoregressive and contingent upon the _entire_ state of the world up until that point.
We use the cross-entropy to maximize the negative log-likelihood of the tokens based upon the training set, this is the best way to directly minimize the empirical risk, at least mathematically speaking.
While some of the information of the world state is inherently unknowable to some degree (i.e., 'noise'), building an understanding of the connections between concepts offers a learned prior that matches the density of the generating distribution (i.e., real life).
Couple this with a severe L2 penalty on the weights, which optimizes for the MDL in the limit (!!!!), and you have a system that fundamentally embeds an approximation of the information graph of the world in some neural network. This is quite literally the _only_ way to improve next token prediction once you get beyond the initial token-occurrence statistics, etc.
In the limit, the only way to reliably predict the world state as accurately as possible without having direct info of the world state at that time is to learn the entire conceptual graph of the world, thus minimizing our achievable log likelihood with the available information that we've been given. _This_ is what Hinton is talking about, as best as I understand.
The author uses a bit of an illusion of shortcutting -- which is an ideal strategy for an _online_ agent with limited resources in a dynamic world, and for models earlier on in their training process. But Hinton is not talking about this at all, really, no! He is talking about the limit! Of course, if you stop an LLM in the middle of training (or look at earlier, smaller ones), you'll see similar 'shortcut' methods. This is a matter of capacity, which is tangentially in the same family as the author's casual, more personal examples, but not at all really related to the mathematics of what is going on behind the scenes here. These are two entirely different problem sets, it's apples to oranges, and there's not really much tie here. From their profile page, the author is a professor of statistics and political science, and I'm not sure why the information theory side of things didn't come up given the statistics background (though they may be somewhat disjoint).
Hinton was being polite to the general public in not dropping all of the math on the reader at once, and I respect that. I understand how someone might misunderstand that and go long on an unrelated rebuttal, but it is frustrating to not see a healthy level of rigor applied here.
I do feel somewhat bothered this is also being upvoted on HN. I know not everyone is a practitioner, but I think this article misses the quality bar. We really gotta just emphasize, and re-emphasize the fundamentals over and over. I feel we may flounder and go on silly tangents otherwise.
Happy to answer any technical questions in the comments.