5 ms·
Useful comment. But I wonder if you are not overstating the anti-meaning interpretation. Yes, attention is task-oriented, but it is instrumentally useful to tra
by ybell 3y ago
Useful comment. But I wonder if you are not overstating the anti-meaning interpretation. Yes, attention is task-oriented, but it is instrumentally useful to track inputs in a way that approximates meaning. Isn't that why we say that initial layers are syntax oriented while later layers are more abstract?
- light_hue_1 3y agoThe goal of attention is simply to find the parts of the input that are useful for the task being carried out. There's no component dedicated to "meaning", there's no objective function for that, etc. The reason why we say initial layers contain information about syntax is because we attach probes (linear decoders) to those layers and run an experiment: can I classify the part of speech of words with the activity at the first layer, at the second layer, etc. Turns out you can use early layers to classify the part of speech of words. But for something like coreference resolution you need much later layers. But that's an emergent property. Nothing in the network talks about meaning in any way. And certainly not about syntax, etc. Definitely doesn't mean those early layers are only about syntax, or about syntax as we think of it: they just have enough information to classify things like part of speech. I'm not being picky. This is really important! That's why networks like this can take as input language, audio, images, video, etc. Because they don't commit to one idea of what meaning means, they just attempt to do masked word/region prediction.