4 ms·
"The second token, however, duplicates the input of the first one and does not receive a gradient "answer" from the very next token, only from future tokens; ..
by wantsanagent 2y ago
"The second token, however, duplicates the input of the first one and does not receive a gradient "answer" from the very next token, only from future tokens; ..."
This formulation doesn't make a lot of sense to me.
I get the motivation here but what you're trying to implement is a working memory.
Because transformers have perfect retrospective memory within their context window any generation which can be done directly from input tokens will be.
At any given point a model might want to write to a working memory, but that does not imply that the next non-working-memory-step will supply useful information to better write to working memory in the future. The model also has to be able to decide when to compare the work done in working memory to the next token.
By allowing the model to both exempt output from gradient updates and opt back in to gradient updates, you create a meta-learning loop that could be quite flexible.
- sdwr 2y agoAs I understand it, this isn't trying to implement actual memory in the form of a cache, but instead some kind of wishy-washy memory-lite. I'm talking out of my ass here, but I feel like real memory shouldn't be that hard to implement on top of chatGPT. Just run it twice per query, the first time as an internal query that fetches from a memory store. The budgeting part would be interesting. How many tokens of the main query do you want to fill with memories? And it wouldn't be able to meta learn how to use the system better, you'd have to update the prompt
- refulgentis 2y agoI'll call it "not even wrong" :P here, they're putting it in the model, you're describing a common bit of working with LLMs across memory / RAG / etc.