6 ms·
The main thesis here seems to be that LLMs behave like almost all other machine learning models, in that they are doing pattern matching on their input data, an
by moolimon 2y ago
The main thesis here seems to be that LLMs behave like almost all other machine learning models, in that they are doing pattern matching on their input data, and short circuiting to a statistically likely result. Chain of thought reasoning is still bound by this basic property of reflexive pattern matching, except the LLM is forced to go through a process of iteratively refining the domain it does matching on.
Chain of thought is interesting, because you can combine it with reinforcement learning to get models to solve (seemingly) arbitrarily hard problems. This comes with the caveat that you need some reward model for all RL. This means you need a clear definition of success, and some way of rewarding being closer to success, to actually solve those problems.
Framing transformer based models as pattern matchers makes all the sense in the world. Pattern matching is obviously vital to human problem solving skills too. Interesting to think about what structures human intelligence has that these models don't. For one, humans can integrate absolutely gargantuan amounts of information extremely efficiently.
- drakenot 2y agoWith DeepSeek-R1-Zero, their usage of RL didn't have reward functions really that indicated progress towards the goal afaik. It was "correct structure, wrong answer", "correct answer", "wrong answer". This was for Math & Coding, where they could verify answers deterministically.
- deleted 2y ago[deleted]
- mountainriver 2y agoIt is a reward function it’s just a deterministic one. Reward models are often hacked preventing real reasoning from being discovered
- huijzer 2y ago> Framing transformer based models as pattern matchers makes all the sense in the world. Pattern matching is obviously vital to human problem solving skills too. Interesting to think about what structures human intelligence has that these models don't. For one, humans can integrate absolutely gargantuan amounts of information extremely efficiently. What is also a benefit for humans, I think, is that people are typically much more selective. LLMs train to predict anything on the internet, so for example for finance that includes clickbait articles which have a lifetime of about 2 hours. Experts would probably reject any information in these articles and instead try to focus on high quality sources only. Similarly, a math researcher will probably have read a completely set of sources throughout the life than, say, a lawyer. I’m not sure it’s a fundamental difference, but current models do seem to not specialize from the start unlike humans. And that might be in the way of learning the best representations. I know from ice hockey for example, that you can see within 3 seconds whether someone played ice hockey from young age or not. Same with language. People can usually hear an accent within seconds. Relatedly, I've used OpenAI's text to speech a while back and the Dutch voice had an American accent. What this means is that even if you ask LLMs about Buffett's strategy, maybe they have a "clickbait accent" too. So with the current approach to training, the models might never reach absolute expert performance.
- andai 2y agoWhen I was doing some NLP stuff a few years ago, I downloaded a few blobs of Common Crawl data, i.e. the kind of thing GPT was trained on. I was sort of horrified by the subject matter and quality: spam, advertisements, flame wars, porn... and that seems to be the vast majority of internet content. (If you've talked to a model without RLHF like one of the base Llama models, you may notice the personality is... different!) I also started wondering about the utility of spending most of the network memorizing infinite trivia (even excluding most of the content above, which is trash), when LLMs don't really excel at that anyway, and they need to Google it anyway to give you a source. (Aside: I've heard soke people have good luck with "hallucinate then verify" with RAG / Googling...) i.e. what if we put those neurons to better use? Then I found the Phi-1 paper, which did exactly that. Instead of training the model on slop, they trained it on textbooks! And instead of starting with PhD level stuff, they started with kid level stuff and gradually increased the difficulty. What will we think of next...
- dr_dshiv 2y agoYes, but the PHI-1 textbooks were synthetic — written by other models! So…
- astrange 2y agoYou can get rid of the trivia by training one model on the slop, then a second model on the first one - called distillation or teacher-student training. But it's not much of a problem because regularization during training should discourage it from learning random noise. The reason LLMs work isn't because they learn the whole internet, it's because they try to learn it but then fail to, in a useful way. If anything current models are overly optimized away from this; I get the feeling they mostly want to tell you things from Wikipedia. You don't get a lot of answers that look like they came from a book.
- gf000 2y agoI don't know, babies hear a lot of widely generic topics from multiple people before learning to speak. I would rather put it that humans can additionally specialize much more, but we usually have a pretty okay generic understanding/model of a thing we consider as 'known'. I would even wager that being generic enough (ergo, has been sufficiently abstracted) is possibly the most important "feature" human's have? (In the context of learning)
- mnky9800n 2y agoHumans often do not have a clear definition of success and instead create a post-hoc narrative to describe whatever happened as success.
- cadamsdotcom 2y agoContinuous RL in a sense. There maybe an undiscovered additional scaling law around models doing what you describe; continuous LLM-as-self-judge, if you will. Provided it can be determined why a user ended the chat, which may turn out to be possible in some subset of conversations.
- ahartmetz 2y agoI'm not following. Do you have an example?
- mnky9800n 2y agoThe milliken oil drop experiment, “winning “ the space race, mostly anything C levels will tell the board and shareholders at a shareholder meeting, the American wars in Iraq and Afghanistan, most of what Sam Altman or Elon musk has to say, this list continues.
- Yiin 2y agoI think you're approaching it form very high level, when you should think about it from much lower level, i.e. success is being determined by stress/dopamine hormones or similar
- d0mine 2y agoThe lower level seems to work eg, “Dopamine regulates decision thresholds in human reinforcement learning” https://www.nature.com/articles/s41467-023-41130-y https://www.nature.com/articles/s41467-023-41130-y
- 2y ago
- mdp2021 2y ago> Interesting to think about what structures human intelligence has that these models don't Chiefly? After having thought long and hard, building further knowledge on the results of the process of having thought long and hard, and creating intellectual keys to further think long and hard better.
- nonameiguess 2y agoTo me: LLMs are trained, as others have mentioned, first to just learn the language at all costs. Ingest any and all strings of text generated by humans until you can learn how to generate text in a way that is indistinguishable. As a happy side effect, this language you've now learned happens to embed quite a few statements of fact and examples of high-quality logical reasoning, but crucially, the language itself isn't a representation of reality or of good reasoning. It isn't meant to be. It's a way to store and communicate arbitrary ideas, which may be wrong or bad or both. Thus, the problem for these researchers now becomes how do we tease out and surface the parts of the model that can produce factually accurate and reasonable statements and dampen everything else? Animal learning isn't like this. We don't require language at all to represent and reason about reality. We have multimodal sensory experience and direct interaction with the physical world, not just recorded images or writing about the world, from the beginning. Whatever it is humans do, I think we at least innately understand that language isn't truth or reason. It's just a way to encode arbitrary information. Some way or another, we all grok that there is a hierarchy of evidence or even what evidence is and isn't in the first place. Going into the backyard to find where your dog left the ball or reading a physics textbook is fundamentally a different form of learning than reading the Odyssey or the published manifesto of a mass murderer. We're still "learning" in the sense that our brains now contain more information than they did before, but we know some of these things are representations of reality and some are not. We have access to the world beyond the shadows in the cave.
- anon84873628 2y agoHumans can carve the world up into domains with a fixed set of rules and then do symbolic reasoning within it. LLMs can't see to do this in a formal way at all -- they just occasionally get it right when the domain happens to be encoded in their language learning. You can't feed an LLM a formal language grammar (e.g. SQL) then have it only generate results with valid syntax. It's awfully confusing to me that people think current LLMs (or multi-modal models etc) are "close" to AGI (for whatever various definitions of all those words you want to use) when they can't do real symbolic reasoning. Though I'm not an expert and happy to be corrected...
- 2y ago
- brazzy 2y ago> Interesting to think about what structures human intelligence has that these models don't. Constant direct feedback from the real world and the ability to continuously integrate it to update the model. That's probably the big one. My pet theory is that having a body is actually an integral part of intelligence, to provide the above, as well as an anchor for a sense of self
- mdp2021 2y ago> having a body You do not need sensorial feedback to do math. And you do not need full sensors to have feeback - one well organized channel can suffice for some applications.
- buovjaga 2y ago
- imtringued 2y agoComing up with a reward model seems to be really easy though. Every decidable problem can be used as reward model. The only downside to this is that the LLM community has developed a severe disdain for making LLMs perform anything that can be verified by a classical algorithm. Only the most random data from the internet will do!
- mnky9800n 2y agoI feel like if you take the underlying transformer and apply to other topics, e.g., eqtransformer, nobody questions this assumption. It’s only when language is in the mix do people suggest they are something more and some kind of “artificial intelligence” akin to the beginnings of Data from Star Trek or C3P0 from Star Wars.
- ben_w 2y ago> For one, humans can integrate absolutely gargantuan amounts of information extremely efficiently. What we can integrate, we seem to integrate efficiently*; but compared to the quantities used to train AI, we humans may as well be literally vegetables. * though people do argue about exactly how much input we get from vision etc., personally I doubt vision input is important to general human intelligence, because if it was then people born blind would have intellectual development difficulties that I've never heard suggested exist — David Blunket's success says human intelligence isn't just fine-tuning on top of a massive vision-grounded model.
- Retric 2y agoHearing is also well into the terabytes worth of information per year. Add in touch, taste, smell, proprioception, etc and the brain gets a deluge. The difference is we’re really focused on moving around in 3D space and more abstract work, where an LLM etc is optimized for a very narrow domain.
- jdietrich 2y ago>Hearing is also well into the terabytes worth of information per year. If we assume that the human auditory system is equivalent to uncompressed digital recording, sure. Actual neural coding is much more efficient, so the amount of data that is meaningfully processed after multiple stages of filtering and compression is plausibly on the order of tens of gigabytes per year; the amount actually retained is plausibly in the tens of megabytes. Don't get me wrong, the human brain is hugely impressive, but we're heavily reliant on very lossy sensory mechanisms. A few rounds of Kim's Game will powerfully reveal just how much of what we perceive is instantly discarded, even when we're paying close attention.
- Retric 2y agoThe sensory information form individual hairs in the ear start off with a lot more data to process than simple digital encoding of two audio streams. Neural encoding isn’t particularly efficient from a pure data standpoint just an energy standpoint. A given neuron not firing is information and those nerve bundles contain a lot of neurons.
- arkh 2y ago> Interesting to think about what structures human intelligence has that these models don't. Pain receptors. If you want to mimic human psyche you have to make your agent want to gather resources and reproduce. And make it painful to lack those resources. Now, do we really have to mimic human intelligence to get intelligence? You could make the point the internet is now a living organism but does it have some intellect or is it just some human parasite / symbiote?
- spenrose 2y agoI argue that we should start calling them "pattern processors": https://x.com/sampenrose/status/1877200883613659360 https://x.com/sampenrose/status/1877200883613659360
- PaulDavisThe1st 2y agoYour post on Twitter uses slightly more words than the ones preceding it above to make the exact same point. Was there really any reason to link to it? Why not expand on your argument here?
- lubujackson 2y agoHuman processing is very interesting and should likely lead to more improvements (and more understanding of human thought!) Seems to me humans are very good at pattern matching, as a core requirement for intelligence. Not only that, we are wired to enjoy it innately - see sudoku, find Waldo, etc. We also massively distill input information into short summaries. This is easy to see by what humans are blind to: the guy in a gorilla suit walking through a bunch of people passing a ball around, or basically any human behavior magicians use to deceive or redirect attention. We are mombarded with information constantly. This is the biggest difference between us and LLMs as we have a lot more input data and also are constantly updating that information - with the added feature/limitation of time decay. It would be hard to navigate life without short term memory or a clear way to distinguish things that happened 10 minutes ago from 10 months ago. We don't fully recall each memory of washing the dishes but junk the vast, vast majority of our memories, which is probably the biggest shortcut our brains have over LLMs. Then we also, crucially, store these summaries in memory as connected vignettes. And our memory is faulty but also quite rich for how "lossy" it must be. Think of a memory involving a ball from before the age of 10 and most people can drum up several relevant memories without much effort, no matter their age.
- viccis 2y ago>Interesting to think about what structures human intelligence has that these models don't. Kant's Critique of Pure Reason has been a very influential way of examining this kind of epistemology. He put forth the argument that our ability to reason about objects comes through our apprehension of sensory input over time, schematizing these into an understanding of the objects, and finally, through reason (by way of the categories) into synthetic a priori knowledge (conclusions grounded in reason rather than empiricism). If we look at this question in that sense, LLMs are good at symbolic manipulation that mimics our sensibility, as well as combining different encounters with concepts into an understanding of what those objects are relative to other sensed objects. What it lacks is the transcendental reasoning that can form novel and well grounded conclusions. Such a system that could do this might consist of an LLM layer for translating sensory input (in LLM's case, language) into a representation that can be used by a logical system (of the kind that was popular in AI's first big boom) and then fed back out.
- corimaith 2y ago>Such a system that could do this might consist of an LLM layer for translating sensory input (in LLM's case, language) into a representation that can be used by a logical system (of the kind that was popular in AI's first big boom) and then fed back out. This just goes back into the problems of that AI winter again though. First Order Logic isn't expressive enough to model the real world, while Second Order Logic dosen't have a complete proof system to truly verify all it'sstatements, and is too complex and unyieldy for practical uses. The number of people I would also imagine that are working on such problems would be very few, this isn't engineering that it is analytic philosophy and mathematics.
- viccis 2y agoKant predates analytical philosophy and some of its failures (the logical positivism you are referring to). The idea here is that first order logic doesn't need to be expressive enough to model the world. Only that some logic system is capable of modeling the understanding of a representation of the world mediated by way of perception (via the current multimodal generative AI models). And finally, it does not need to be complete or correct, just equivalent or better than how our minds do such.
- corimaith 2y ago>Interesting to think about what structures human intelligence has that these models don't. If we get to the gritty details of what gradient descent is doing, we've got a "frame", i.e a matrix or some array of weights contains the possible solution for a problem, then with another input of weights we're matching a probability distribution to minimize the loss function with our training data to form our solution in the "frame". That works for something like image recognition, where the "frame" is just the matrix of pixels, or in language models where we're trying to find the next word-vector given a preceding input. But take something like what Sir William Rowan Hamilton was doing back in 1843. He know that complex numbers could be represented in points in a plane, and arthimetic could be performed on them, and now he wanted to extend a similar way for points in a space. With triples it is easy to define addition, but the problem was multiplication. In the end, he made an intuitive jump, a pattern recognition when he realized that he could easily define multiplications used quadruples instead, and thus was born the Quaternion that's a staple in 3D graphics today. If we want to generalize this kind of problem solving into a way that gradient descent can solve, where do we even start? First of all, we don't even know if a solution is possible or coherent or what "direction" we are going towards. It's not a systematic solution, it's rather one that pattern in one branch of mathematics was recognized into another. So perhaps you might use something like Category Theory, but then how are we going to represent this in terms of numbers and convex functions, and is Category Theory even practical enough to easily do this?
- 1vuio0pswjnm7 2y ago"LLMs are fundamentally matching the patterns they've seen, and their abilities are constrained by mathematical boundaries. Embedding tricks and chain-of-thought prompting simply extends their ability to do more sophisticated pattern matching."