3 ms·
You haven't provided any counter argument, and there are many articles with data backing this up. Unfortunately I am not aware of any articles that can show co
by chaxor 3y ago
You haven't provided any counter argument, and there are many articles with data backing this up.
Unfortunately I am not aware of any articles that can show convincing data that "LLMs don't learn like humans". I'm not really even sure what that means precisely.
Your statement could be understood to claim that language models cannot predict brain activations or vice versa? Predictive power is the foundation of science, so it's one of the better tests we have for these types of problems. To that point, there is substantial evidence of similarity and predictive power, as noted below.
Perhaps you mean something else? Surely you're point is not that GPUs simply aren't human. Of course - there are differences at various levels. So at some level, everyone recognizes they're not identical.
Perhaps it's that attention architectures are not reproduced within neural microarchitectures in tbe neocortex? I haven't seen any studies on this within the relevant cortical areas to address this, though microarchitecture matching may not be required for output distribution matching.
The interesting endeavour for science to show right now is how these systems are similar to our processing mechanisms. The ways in which they're not similar typically tend to be uninteresting, or obvious.
The fact that we can predict brain activity from NN model activations raises many more opportunities in science, and the predictive power has gotten far better with these advances,.
Here are a small set of articles from MIT, deepmind, etc. for some of the points I made:
- Tuckute, Greta, Aalok Sathe, Shashank Srikant, Maya Taliaferro, Mingye Wang, Martin Schrimpf, Kendrick Kay, and Evelina Fedorenko. "Driving and suppressing the human language network using large language models." Nature Human Behaviour (2024): 1-18.
- Schrimpf, Martin, Jonas Kubilius, Ha Hong, Najib J. Majaj, Rishi Rajalingham, Elias B. Issa, Kohitij Kar et al. "Brain-score: Which artificial neural network for object recognition is most brain-like?." BioRxiv (2018): 407007.
- Pereira, Francisco, Bin Lou, Brianna Pritchett, Samuel Ritter, Samuel J. Gershman, Nancy Kanwisher, Matthew Botvinick, and Evelina Fedorenko. "Toward a universal decoder of linguistic meaning from brain activation." Nature communications 9, no. 1 (2018): 963.
- Arend, Luke, Yena Han, Martin Schrimpf, Pouya Bashivan, Kohitij Kar, Tomaso Poggio, James J. DiCarlo, and Xavier Boix. Single units in a deep neural network functionally correspond with neurons in the brain: preliminary results. Center for Brains, Minds and Machines (CBMM), 2018.
- foobarqux 3y agoI did provide a counter argument and referenced evidence "backing it up": Namely LLMs can learn languages that are not human-compatible (Moro, "Secrets of Words"). I seriously doubt you have read, never mind understood, any of the papers you cite; if you had we could discuss what's wrong with them. Without even having read them there are obvious problems like the fact that you can't measure a large number of individual neuron activation in the brain and that different deep learning networks have substantially different activations so they cannot be meaningfully be similar to the brain if they aren’t even similar to one another (or that the similarity is so nebulous as to be meaningless). But this is moot because as I said at the outset there are fundamental differences at the functional level which make existing LLMs (and deep learning) unlike brains.
- chaxor 3y agoI think we may agree more than you realize, but perhaps we aren't clear about what we mean. As I said pretty early on - they are similar at some level. To make an analogy, that could be as simple as a submarine (w2v) and a fish (authentic language) now 'uses fins' to move (bert). They're now more similar, but not the same. Of course there are differences, but we can learn by studying what structure has been added and how in order to work towards better theory (as well as looking at what differences there are). I have read all of these studies, as I am an academic in related fields, and you are correct that the resolution is not as good as we would hope (it's better in the ventral visual stream in non human primates, which is one reason we started there); however it doesn't negate the importance of the predictive power here. We can make very useful tools out of these models which can improve the lives of non-verbal people for example. Another interesting set of works going in recently is in testing out the poverty of stimulus. There is now evidence from many labs that LMs learn with the same amount of data as humans, so the sample efficiency arguments and the 'poverty of stimulus' aren't really as successful as arguments for training difference either. To the point of humans not learning certain languages, I haven't seen it proven that humans are absolutely incapable of learning certain languages. I do know that children ignore much of sequence structure and create their own internal structures, but this does not negate learning languages which have some arbitrary set of rules. I'm aware that models learn DNA more easily _practically_, but this is more of a statement about human experience and various other environmental factors - not a mathematical proof that it's impossible for humans to learn. Furthermore, that is a nice property of these systems. If they are capable of modeling many different phenomenon, they are useful for many things. If they can settle on a model configuration that linearly maps to a similar activation space with similar topological properties to a patient, that's also very useful. In other words, If the topological properties of the activation space are more constrained in the cortex vs more flexible NNs, but the NNs can fit to the constrained space, that doesn't seem like an insurmountable problem in studying either trained system. Ultimately, I think we can agree that there are differences, but there are also striking similarities that can be very useful for improving science and medicine going forward, and perhaps (with _substantial_ effort in linguistics and mechanistic interpretability fields) we may be able to improve some of our understanding of linguistics and/or neuroscience (or perhaps not, but it's at least a promising potential lead).