4 ms·
I would argue the holdup right now is long term memory. GPTs already have the ability to rapidly generalize and incorporate new knowledge within the context win
by valine 3y ago
I would argue the holdup right now is long term memory. GPTs already have the ability to rapidly generalize and incorporate new knowledge within the context window. The trick is to retain what it has learned.
It won’t take a very long time to fix that.
This is model I trained with a fine tuning technique based on this idea. The training dataset consists of instructions like “Talk like a pirate”. The concept generalized well and the model responds in the style of a pirate far more consistently than an equivalent system prompt.
https://huggingface.co/valine/OpenPirate https://huggingface.co/valine/OpenPirate
Offloading context learning into the model weights frees you from the computation and memory burden of the attention mechanism. I expect a technique like this will probably be a piece of AGI someday.
- byyoung3 3y agoJust need a few backward passes and u get long term memory. I think we are overthinking that aspect
- valine 3y agoTakes way more than that in my experience. Back prop isn’t sufficient for rapid generalization.
- byyoung3 3y agoWell, I think you definitely need a certain level of scale, but I think it definitely still works. Also, generalization and memory are two very different things. Generalization is basically iQ, which is very much still a work in progress.
- valine 3y agoIf you have rapid generalization you don’t need scale. Large datasets are only necessary to compensate for the lack of good generalization. The model I posted in my earlier comment responds in character for all queries and was trained in 60 seconds with a dataset smaller than this comment thread.
- byyoung3 3y ago"If you have rapid generalization you don’t need scale" As of right now, you do need scale for rapid generalization. The technology for NLP generalization without scale (both model and data) does not exist. Not saying it won't in the future, just not right now
- valine 3y agoIt does exist, I’ve been playing around with it all week :). Let me know if you’d like a custom Mistral 7B I’ll train one for you. From the output of bing chat, and the very specific way it goes off the rails, I suspect Microsoft has figured it out too. The algorithmic jump to rapid generalization isn’t hard to make. I would be shocked if there’s not an open source version of it a year from now.