5 ms·
Does this mean it is now computationally efficient to have the model learn/memorize information on the fly, say the current chat context, as part of the model w
by SubiculumCode 3y ago
Does this mean it is now computationally efficient to have the model learn/memorize information on the fly, say the current chat context, as part of the model weights? One shot encoding (something the hippocampus is very good at) allows us to build experiences into retrievable memories tied into semantic concepts we've previously learned..in fact it gets better the more rich our semantic conceptualization of events become from childhood into adulthood.
If memorization of events in llm is accelerated because of- these deep semantic frameworks, then does this provide a path towards long context windows?
- quickthrower2 3y agoBeginner here, so just musing: I like the idea. You would need your own mutable copy of the model, which is usually huge. And you need to backprop so there is a bit more computation. It might be doable for a local model that is smaller than GPT3.5/4. You also need to decide what is worth memorizing long term vs short term.
- pests 3y ago> own mutable copy of the model, which is usually huge It could just be the diff against the main model or similar.
- quickthrower2 3y agoBut if you have say 50bn weights, and you run backprop, you are going to update most of the weights (except the dropout ones, but which ones drop out changes on every token I think). This means you need 50bn deltas. It might compress, but if you do then you need extra compute to do that.
- jacquesm 3y agoYou would do dropout on every epoch of training, not on every token.
- quickthrower2 3y agoI didn't know that, I might look at the NanoGPT code and torch.dropout docs a bit closer then. Thanks!
- SubiculumCode 3y agoComing back to this. LORA training is only on the attention layer, and this was sufficient for memorization , per the article. So we wouldn't update all the model's weights in some kind of constant context one-shot learning scheme.
- warkdarrior 3y agoMaybe, but there are a lot of unknowns. Does the "memorization on the fly" come with catastrophic forgetting of other information? How does one control for memorizing recent stuff vs. remembering older stuff?