5 ms·
Can LLMs take advantage of this bigger window to solve meaningful tasks though? I can't imagine in the training data, knowing what happened 100k tokens ago woul
by sweezyjeezy 3y ago
Can LLMs take advantage of this bigger window to solve meaningful tasks though? I can't imagine in the training data, knowing what happened 100k tokens ago would be _that_ relevant to predicting the current token very often, so unless this is something that the model learns to leverage more implicitly, I'd be a bit pessimistic.
- ttul 3y agoYes. For instance, a large context window allows you to have a chat for months where the model can remember and make use of everything you’ve ever talked about. That enables creating a much more effective “assistant” that can remember key details months later that may be valuable. A second example is the analysis of long documents. Today, hacks like chunking and HyDE enable us to ask questions about a long document or a corpus of documents. But is far superior if the model can ingest the whole document and apply attention to everything, rather than just one chunk at a time. Chunking effectively means that the model is limited to drawing conclusions from one chunk at a time and cannot synthesize useful responses relating to the entire document.
- sweezyjeezy 3y agoI'm not questioning whether it would be useful, just whether it's actually something that token masking in training is going to work to make the model learn this.
- kaj_sotala 3y agoI would imagine that if you are training on the text of a novel, then anything that happened earlier in the text may be relevant for predicting the next events. Especially if it's something like a detective novel that has clues about the criminal's identity scattered across the story. Also if you are training on a database of code.
- sweezyjeezy 3y agoYeah but when you're training a neural net with backprop on a finite dataset, "this would help the model" ≠ "the model will learn this". This is 100% speculation, but my intuition is that it's not going to work very well unless it happens 'a lot' in the training data, or if they've curated the data specifically to try and make it learn long range signals.
- woeirua 3y agoIt remains to be seen just how effective longer contexts are because if the attention vectors don't ever learn to pick up specific items from further back in the text then having more tokens doesn't really matter. Given that the conventional cost of training attention layers grows quadratically with the number of tokens I think Anthropic is doing some kind of approximation here. Not clear at all that you would get the same results as vanilla attention.
- ttul 3y agoThey did mention that the inference time to answer a question about the book was something like 22 seconds, so perhaps they are indeed still using self-attention.
- m3kw9 3y agoGets pricier as you chat for longer, imagine having to chat a line with a history with 20k token.
- SomewhatLikely 3y agoI would guess that semantic similarity would be the stronger training signal than distance once you go beyond a sentence or two away.
- sweezyjeezy 3y agoI'm pretty dubious - how would the model not get absolutely swamped by the vast amount of potential context if it's not learning to ignore long range signals for the most part?
- redox99 3y agoI'd argue that books are a clear example where the 100k tokens context would make a huge difference.