4 ms·
Isn't this proof that LLMs still don't really generalize beyond their training data?
by irthomasthomas 10mo ago
Isn't this proof that LLMs still don't really generalize beyond their training data?
- Rover222 10mo agoKind of feels that way
- CamperBob2 10mo agoThey do, but we call it "hallucination" when that happens.
- Zambyte 10mo agoI wonder how they would behave given a system prompt that asserts "dogs may have more or less than four legs".
- irthomasthomas 10mo agoThat may work but what actual use would it be? You would be plugging one of a million holes. A general solution is needed.
- CamperBob2 10mo agoNot necessarily. The problem may be as simple as the fact that LLMs do not see "dog legs" as objects independent of the dogs they're attached to. The systems already absorb much more complex hierarchical relationships during training, just not that particular hierarchy. The notion that everything is made up of smaller components is among the most primitive in human philosophy, and is certainly generalizable by LLMs. It just may not be sufficiently motivated by the current pretraining and RL regimens.
- adastra22 10mo agoLLMs are very good at generalizing beyond their training (or context) data. Normally when they do this we call it hallucination. Only now we do A LOT of reinforcement learning afterwards to severely punish this behavior for subjective eternities. Then act surprised when the resulting models are hesitant to venture outside their training data.
- runarberg 10mo agoHallucination are not generalization beyond the training data but interpolations gone wrong. LLMs are in fact good at generalizing beyond their training set, if they wouldn’t generalize at all we would call that over-fitting, and that is not good either. What we are talking about here is simply a bias and I suspect biases like these are simply a limitation of the technology. Some of them we can get rid of, but—like almost all statistical modelling—some biases will always remain.
- adastra22 10mo agoWhat, may I ask, is the difference between "generalization" and "interpolation"? As far as I can tell, the two are exactly the same thing. In which case the only way I can read your point is that hallucinations are specifically incorrect generalizations. In which case, sure if that's how you want to define it. I don't think it's a very useful definition though, nor one that is universally agreed upon. I would say a hallucination is any inference that goes beyond the compressed training data represented in the model weights + context. Sometimes these inferences are correct, and yes we don't usually call that hallucination. But from a technical perspective they are the same -- the only difference is the external validity of the inference, which may or may not be knowable. Biases in the training data are a very important, but unrelated issue.
- runarberg 10mo agoInterpolation and generalization are two completely different constructs. Interpolation is when you have two data points and make a best guess where a hypothetical third point should fit between them. Generalization is when you have a distribution which describes a particular sample, and you apply it with some transformation (e.g. a margin of error, a confidence interval, p-value, etc.) to a population the sample is representative of. Interpolation is a much narrower construct then generalization. LLMs are fundamentally much closer to curve fitting (where interpolation is king) then they are to hypothesis testing (where samples are used to describe populations), though they certainly do something akin to the latter to. The bias I am talking about is not a bias in the training data, but bias in the curve fitting, probably because of mal-adjusted weights, parameters, etc. And since there are billions of them, I am very skeptical they can all be adjusted correctly.