4 ms·
I wonder if you could perform inference, highlight the weights that were most used during that inference, grab the hottest 20%, freeze the rest of the model, an
by armatav 3y ago
I wonder if you could perform inference, highlight the weights that were most used during that inference, grab the hottest 20%, freeze the rest of the model, and perform backpropagation solely on those to allow for more of this sort of rapid memorization behavior closer to the end user.
Like online learning in a way. But you do it during inference time.
There’s no way the entire model actually needs to be touched for something like “sky color is:” and “blue”.
- armatav 3y agoIn fact I bet you could update like one or two neurons for certain concepts, and then transplant those neurons to another LLM to give it some idea of it. Like a literal brain transplant but for concepts.
- armatav 3y agoAnd you could identify these neurons using dropout techniques and repetitively querying the model against them. Drop a set of neurons and there’s no change? Probably doesn’t contain the “sky color” concept. Drop a set of neurons and the model freaks out, definitely conceptual neurons. Rinse and repeat to find the distilled pattern across all the neurons. You could train an LLM against the neuron graph to do this for you.
- niemandhier 3y agoMany neurons are polysynthactic, that makes interventions like the proposed difficult.
- armatav 3y agoIs that necessarily the case for massive LLMs, or could there be a more refined grouping?
- niemandhier 3y agoI am not sure if the origin of polysyntacicity is fully understood. From a physics perspective it’s entropy: There are just more local minima that have neurons code multiple things. I suspect that dropout and similar tricks also play a role in this: Removing connections during training means that pathways need to be redundant somewhat.
- armatav 3y agoThat's super interesting, I wonder if dropout applied in a certain pattern of segregation could induce more neuron dependence, instead of less. Or do some anti-dropout where a dedicated portion of the model is used. Not sure why you'd want that, but it might be interesting to explore.