4 ms·
That's a fair point, multiple models is a potential implementation detail for parallelization - I do think it's fair to assume you are "restarting" the model or
by beardedwizard 2y ago
That's a fair point, multiple models is a potential implementation detail for parallelization - I do think it's fair to assume you are "restarting" the model or clearing context by some means between runs otherwise I don't think it achieves the goal of groomed attention.
- mewpmewp2 2y agoThat gets really philosophical here. What is a single run actually? Because each token is generated one by one, with all the previous tokens as input. If we remove one token from the context in the past does it make it a new "run"? What if we have a summarization memory system, where in order to keep the context size small after certain length it will start to summarize/compress the beginning until the context size is good. Then the whole input and context is constantly changing/evolving. You could have 100s of different randomly selected instances generating a new token one by one to the (last input + token generated by another instance) as input.
- beardedwizard 2y agoI think it comes down to the goal you are trying to achieve: maximize use of available attention while minimizing drift. The experiment seems clear, and the outcome being tested is the resulting generation.