5 ms·
Unfortunately its not very good at longer context lengths, which sort of defeats the point of efficient scaling with context. See https://twitter.com/arankomats
by Straw 4y ago
Unfortunately its not very good at longer context lengths, which sort of defeats the point of efficient scaling with context. See https://twitter.com/arankomatsuzaki/status/1639000379978403853 https://twitter.com/arankomatsuzaki/status/16390003799784038...
Its also not really an RNN. The best way to describe the key time mixing operation is a normalized exponentially weighted moving average (EMA)- no non-linearity. Once viewed this way, its not surprising that it struggles at longer contexts- everything decays, and it has limited space to put things. Of course, it does have some clever tricks, and can choose to remember things for a while by upweighting them, but not forever.
- Der_Einzige 4y agoYup. This is why GPT-4 long context length version is going to be such a god damn gamechanger. The current memory techniques we have outside of ultra long context lengths are lossy and imperfect. I wish the langchain spammers (yes it's a good tool) would acknowledge this when they keep posting everywhere about the "memory module".
- lysecret 4y agoYes, I am so curious how they did it, I know about flash attention but there is no way this gets us all the way there.
- sebzim4500 4y agoWhy not? They charge 8x more for a 32k context than for a 8k context (note that the prices are normally presented per token, here I'm talking about absolute cost). Naive scaling on the self-attention component would suggest 16x compute and 4x memory, while the rest of it (forward layers, embeddings, activation functions etc.) all would go up 4x both compute and memory.
- zaptrem 4y agoIt already is a game changer, Bing Chat's Creative mode appears to be using >8k token context (though I'm not sure if it's 16k or 32k).