4 ms·
I disagree. The way I imagine LLMs is that we feed huge amounts of text to a ANN. This ANN doesn't have enough weights to just memorize these huge amounts. So i
by mo_42 3y ago
I disagree. The way I imagine LLMs is that we feed huge amounts of text to a ANN. This ANN doesn't have enough weights to just memorize these huge amounts. So it has to find meaningful abstractions from the texts to make the loss function small. Let's even ignore regularization for now.
For example, it will not memorize that all the instances in the context of earth fall on the ground. It will learn a the meaningful abstraction gravitation. From that it infers that objects should also fall on the ground.
This should be very easy to verify. Moreover, I have also seen some interesting examples of spacial reasoning in ChatGPT, which seems much more complicated than my example.
Edit: I think understanding and finding suitable abstractions is the same thing. That's what we do with children and students. They may try to memorize everything we teach them. But they should find abstractions and commonalities that are useful in more than one instance.