4 ms·
> Compared to humans, LLMs have effectively unbounded training data. They are trained on billions of text examples covering countless topics, styles, and domain
by ForceBru 10mo ago
> Compared to humans, LLMs have effectively unbounded training data. They are trained on billions of text examples covering countless topics, styles, and domains. Their exposure is far broader and more uniform than any human's, and not filtered through lived experience or survival needs.
I think it's the other way round: humans have effectively unbounded training data. We can count exactly how much text any given model saw during training. We know exactly how many images or video frames were used to train it, and so on. Can we count the amount of input humans receive?
I can look at my coffee mug from any angle I want, I can feel it in my hands, I can sniff it, lick it and fiddle with it as much as I want. What happens if I move it away from me? Can I turn it this way, can I lift it up? What does it feel like to drink from this cup? What does it feel like when someone else drinks from my cup? The LLM has no idea because it doesn't have access to sensory data and it can't manipulate real-life objects (yet).
- cortesoft 10mo agoNot only that, but humans also have access to all of the "training data" of hundreds of millions of years of evolution baked into our brains.
- ACCount37 10mo agoWhich must be doing some heavy lifting. Humans ship with all the priors evolution has managed to cram into them. LLMs have to rediscover all of it from scratch just by looking at an awful lot of data.
- hathawsh 10mo agoOTOH, all that data is built on patterns that evolved from many years of evolution, so I think the LLM benefits from that evolution also.
- ACCount37 10mo agoSure, but LLMs are trying to build the algorithms of the human mind backwards, converge on similar functionality based on just some of the inputs and outputs. This isn't an efficient or a lossless process. The fact that they can pull it off to this extent was a very surprising finding.
- layer8 10mo agoI don’t think the amount of data is essential here. The human genome is only around 750 MB, much less than current LLMs, and likely only a small fraction of it determines human intelligence. On the other hand, current LLMs contain immense amounts of factual knowledge that a human newborn carries zero information about. Intelligence likely doesn’t require that much data, and it may be more a question of evolutionary chance. After all, human intelligence is largely (if not exclusively) the result of natural selection from random mutations, with a generation count that’s likely smaller than the number of training iterations of LLMs. We haven’t found a way yet to artificially develop a digital equivalent effectively, and the way we are training neural networks might actually be a dead end here.
- ACCount37 10mo agoThat just says "low Kolmogorov complexity". All the priors humans ship with can be represented as a relatively compact algorithm. Which gives us no information on computational complexity of running that algorithm, or on what it does exactly. Only that it's small. LLMs don't get that algorithm, so they have to discover certain things the hard way.
- emp17344 10mo agoIt’s unlikely sensory data contributes to intelligence in human beings. Blind people take in far, far less sensory data than sighted people, and yet are no less intelligent. Think of Helen Keller - she was deafblind from an early age, and yet was far more intelligent than the average person. If your hypothesis is correct, and development of human intelligence is primarily driven by sensory data, how do you reconcile this with our observations of people with sensory impairments?
- jakeinspace 10mo agoBlind people tend to have less spatial intelligence though, like significantly more. Not very nice to say like that, and of course they often develop heightened intelligence in other areas, but we do consider human-level spatial reasoning a very important goal in AI.
- emp17344 10mo agoPeople with sensory impairments from birth may be restricted in certain areas, on account of the sensory impairment, but are no less generally cognitively capable than the average person.
- erichocean 10mo ago> but are no less generally cognitively capable than the average person I think this would depend entirely on how the sensory impairment came about, since most genetic problems are not isolated, but carry a bunch of other related problems (all of which can impact intelligence). Lose your eye sight in an accident? I would grant there is likely no difference on average. Otherwise, the null hypothesis is that intelligence (and a whole host of other problems) are likely worse, on average.
- dpark 10mo ago> It’s unlikely sensory data contributes to intelligence in human beings. This is clearly untrue. All information a human ever receives is through sensory data. Unless your position is that the intelligence of a brain that was grown in a vat with no inputs would be equivalent to that of a normal person. Now, does rotating a coffee mug and feeling its weight, seeing it from different angles, etc. improve intelligence? Actually, still yes, if your intelligence test happens to include questions like “is this a picture of a mug” or “which of these objects is closest in weight to a mug”.
- moffkalast 10mo agoThere's only so much information content you can get from a mug though. We get a lot of high quality data that's relatively the same. We run the same routines every day, doing more or less the same things, which makes us extremely reliable at what we do but not very worldly. LLMs get the opposite: sparse, relatively low quality, low modality data that's extremely varied, so they have a much wider breadth of knowledge but they're pretty fragile in comparison since they get relatively little experience on each topic and usually no chance to affirm learning with RL.
- ForceBru 10mo agoYep, LLMs have a greater breadth of knowledge, but it's shallow. Humans are able to achieve much greater depth because they have more data about the subject.
- mdahardy 10mo agoThis is a fair criticism we should've addressed. There's actually a nice study on this: Vong et al. (https://www.science.org/doi/10.1126/science.adi1374 https://www.science.org/doi/10.1126/science.adi1374) hooked up a camera to a baby's head so it would get all the input data a baby gets. A model trained on this data learned some things babies do (eg word-object mappings), but not everything. However, this model couldn't actively manipulate the world in the way that a baby does and I think this is a big reason why humans can learn so quickly and efficiently. That said, LLMs are still trained on significantly more data pretty much no matter how you look at it. E.g. a blind child might hear 10-15 million words by age 6 vs. trillions for LLMs.
- JohnFen 10mo ago> hooked up a camera to a baby's head so it would get all the input data a baby gets. A camera hooked up to the baby's head is absolutely not getting all the input data the baby gets. It's not even getting most of it.
- omneity 10mo agoWhile an LLM is trained on trillions of tokens to acquire its capabilities, it does not actively retain or recall the vast majority of it, and often enough is not able to make deductive reasoning either (e.g. X owns Y does not necessarily translate to Y belongs to X). The acquired knowledge is a lot less uniform than you’re proposing and in fact is full of gaps a human would never make. And more critically, it is not able to peer into all of its vast knowledge at once, so with every prompt what you get is closer to an “instance of a human” than “all of humanity” as you might think of LLMs. (I train and dissect LLMs for a living and for fun)
- minraws 10mo agoI think you are proposing something that's orthogonal to the OP's point. They mentioned the training data is much higher for an LLM, LLM's recall not being uniform was never in question. No one expects compression to be without loss when you scale below knowledge entropy that exists in your training set. I am not saying LLMs do simple compression but just pointing a mathematical certainity. (And I think you don't need to be an expert in creating LLMs to understand them, albeit I think a lot of people here have experience with it aswell so I find the additional emphasis on it moot).
- lumost 10mo agoA big challenge is that the LLM cannot selectively sample it's training set. You don't forget what a coffee cup looks like just because you only drank water for a week. LLMs on the other hand will catastrophically forget anything in their training set when the training set does not have a uniform distribution of samples in each batch.