7 ms·
Related to ways of understanding neural networks, I've seen these views expressed a lot, which to me seem like misconceptions: - LLMs are basically just slight
by montebicyclelo 1y ago
Related to ways of understanding neural networks, I've seen these views expressed a lot, which to me seem like misconceptions:
- LLMs are basically just slightly better `n-gram` models
- The idea of "just" predicting the next token, as if next-token-prediction implies a model must be dumb
(I wonder if this [1] popular response to Karpathy's RNN [2] post is partly to blame for people equating language neural nets with n-gram models. The stochastic parrot paper [3] also somewhat equates LLMs and n-gram models, e.g. "although she primarily had n-gram models in mind, the conclusions remain apt and relevant". I guess there was a time where they were more equivalent, before the nets got really really good)
[1] https://nbviewer.org/gist/yoavg/d76121dfde2618422139 https://nbviewer.org/gist/yoavg/d76121dfde2618422139
[2] https://karpathy.github.io/2015/05/21/rnn-effectiveness/ https://karpathy.github.io/2015/05/21/rnn-effectiveness/
[3] https://dl.acm.org/doi/pdf/10.1145/3442188.3445922 https://dl.acm.org/doi/pdf/10.1145/3442188.3445922
- colah3 1y agoI guess I'll plug my hobby horse: The whole discourse of "stochastic parrots" and "do models understand" and so on is deeply unhealthy because it should be scientific questions about mechanism, and people don't have a vocabulary for discussing the range of mechanisms which might exist inside a neural network. So instead we have lots of arguments where people project meaning onto very fuzzy ideas and the argument doesn't ground out to scientific, empirical claims. Our recent paper reverse engineers the computation neural networks use to answer in a number of interesting cases (https://transformer-circuits.pub/2025/attribution-graphs/biology.html https://transformer-circuits.pub/2025/attribution-graphs/bio... ). We find computation that one might informally describe as "multi-step inference", "planning", and so on. I think it's maybe clarifying for this, because it grounds out to very specific empirical claims about mechanism (which we test by intervention experiments). Of course, one can disagree with the informal language we use. I'm happy for people to use whatever language they want! I think in an ideal world, we'd move more towards talking about concrete mechanism, and we need to develop ways to talk about these informally. There was previous discussion of our paper here: https://news.ycombinator.com/item?id=43505748 https://news.ycombinator.com/item?id=43505748
- mdp2021 1y agoAbsolutely, the first task should be to understand how and why black boxes with emergent properties actually work, in order to further knowledge - but importantly, in order to improve them and build on the acquired knowledge to surpass them. That implies, curbing «parrot[ing]» and inadequate «understand[ing]». I.e. those higher concepts are kept in mind as a goal. It is healthy: it keeps the aim alive.
- HarHarVeryFunny 1y ago1) Isn't it unavoidable that a transformer - a sequential multi-layer architecture - is doing multi-step inference ?! 2) There are two aspects to a rhyming poem: a) It is a poem, so must have a fairly high degree of thematic coherence b) It rhymes, so must have end-of-line rhyming words It seems that to learn to predict (hence generate) a rhyming poem, both of these requirements (theme/story continuation+rhyming) would need to be predicted ("planned") at least by the beginning of the line, since they are inter-related. In contrast, a genre like freestyle rap may also rhyme, but flow is what matters and thematic coherence and rhyming may suffer as a result. In learning to predict (hence generate) freestyle, an LLM might therefore be expected to learn that genre-specific improv is what to expect, and that rhyming is of secondary importance, so one might expect less rhyme-based prediction ("planning") at the start of each bar (line).
- somewhereoutth 1y agoRegardless of the mechanism, the foundational 'conceit' of LLMs is that by dumping enough syntax (and only syntax) into a sufficiently complex system, the semantics can be induced to emerge. Quite a stretch, in my opinion (cf. Plato's Cave).
- Nevermark 1y agoAnyone who has widely read topics across philosophy, science (physics, biology), economics, politics (policy, power), from practitioners, from original takes, news, etc. ... has managed to understand a tremendous number of relationships due to just words and their syntax. While many of these relationships are related to things we see and do in trivial ways, the vast majority go far beyond anything that can be seen or felt. What does economics look like? I don't know, but I know as I puzzle out optimums, or expected outcomes, or whatever, I am moving forms around in my head that I am aware of, can recognize and produce, but couldn't describe with any connection to my senses. The same when seeking a proof for a conjecture in an idiosyncratic algebra. Am I really dealing in semantics? Or have I just learned the graph-like latent representation for (statistical or reliable) invariant relationships in a bunch of syntax? Is there a difference? Don't we just learn the syntax of the visual world? Learning abstractions such as density, attachment, purpose, dimensions, sizes, that are not what we actually see, which is lots of dot magnitudes of three kinds. And even those abstractions benefit greatly from the words other people use describing those concepts. Because you really don't "see" them. I would guess that someone who was born without vision, touch, smell or taste, would still develop what we would consider a semantic understanding of the world, just by hearing. Including a non-trivial more-than-syntactic understanding of vision, touch, smell and taste. Despite making up their own internal "qualia" for them. Our senses are just neuron firings. The rest is hierarchies of compression and prediction based on their "syntax".
- agentcoops 1y ago1000%. It's really hard to express this to non-engineers who never wasted years of their life trying to work with n-grams and NLTK (even topic models) to make sense of textual data... Projects I dreamed of circa 2012 are now completely trivial. If you do have that comparison ready-at-hand, the problem of understanding what this mind-blowing leap means, to which end I find writing like the OP helpful, is so fascinating and something completely different than complaining that it's a "black box." I've expressed this on here before, but it feels like the everyday reception of LLMs has been so damaged by the general public having just gotten a basic grasp on the existence of machine learning.