3 ms·
The context window size - if it really works as advertised - is pretty ground-breaking. It would replace the need to RAG or fine tune for one-off (or few-off) a
by ComputerGuru 3y ago
The context window size - if it really works as advertised - is pretty ground-breaking. It would replace the need to RAG or fine tune for one-off (or few-off) analys{is,es} of input streams cheaper and faster. I wonder how they got past the input token stuffing problems everyone else runs into.
- jcuenod 3y agoIt won't remove the use of RAG at all. That's like saying, "wow, now that I've upgraded my 128GB HDD to 1TB, I'll never run out of space again."
- madisonmay 3y agoIt's more like saying "I've upgraded to 128GB of RAM, I'll never use my disk again".
- sebzim4500 3y ago10 TB for an accurate proportion. And I think people who buy a laptop with a 1TB SSD generally don't run out of space, at least I don't.
- lumost 3y agoThey are almost certainly using some form of sparse attention. If you linearize the attention operation, you can scale up to around 1-10M tokens depending on hardware before hitting memory constraints. Linearization works off the assumption that for a subsequence of X tokens out M tokens, where M os much greater than X there are likely only K tokens which are useful for the attention operation. There are a bunch of techniques to do this, but it's unclear how well any of them scale.
- ein0p 3y agoNot "almost", but certainly. Dense attention is quadratic, not even Google would be able to run it at an acceptable speed. Their model is not recurrent - they did not have the time yet (or resources - believe it or not, Google of 2023-24 is very compute constrained) to train newer SSM or recurrent based models at practical parameter counts. Then there's the fact that those models are far harder to train due to instabilities, which is one of the reasons why you don't yet see FOSS recurrent/SSM models that are SOTA at their size or tokens/sec. With sparse attention, however, long context recall will be far from perfect, and the longer the context the worse the recall. That's better than no recall at all (as in a fully dense attention model which will simply lop off the preceding parts of the conversation), but not by a hell of a lot.
- kiraaa 3y agomaybe they are using ring attention, on top of their 128k model.
- ein0p 3y agoMore likely some clever take on RAG. There’s no way that 1M context is all available at all times. More likely parts of it are retrievable on demand. Hence the retrieval-like use cases you see in the demos. The goal is to find a thing, not to find patterns at a distance
- kiraaa 3y agocould be true, we can only speculate.
- popinman322 3y agovs RAG: RAG is good for searching across >billions of tokens and providing up-to-date information to a static model. Even with huge context lengths it's a good idea to submit high quality inputs to prevent the model from going off on tangents, getting stuck on contradictory information, etc.. vs fine tuning: smaller, fine-tuned models can perform better than huge models in a decent number of tasks. Not strictly fine-tuning, but for throughput limited tasks it'll likely still be better to prune a 70B model down to 2B, keeping only the components you need for accurate inference. I can see this model being good for taking huge inputs and compressing them down for smaller models to use.
- nbardy 3y agoRAG will stick around, at some point you want to retrieve grounded information samples to inject in the context window. RAG+long context just gives you more room for grounded context. Think building huge relevant context on topics before answering.
- torginus 3y agoTbh, I haven't read the paper, but I think it's pretty self-evident that large contexts aren't cheap - the AI has to comb through every word of the context for each successive generated token at least once, so it's going to be at least linear.
- Havoc 3y agoSaw testing earlier that suggested the context does indeed work right