7 ms·
The benefit of "traditional" O(N^2) transformer attention is you correlate every token to every other token. So, in the limit, your network won't "miss" much.
by gamegoblin 3y ago
The benefit of "traditional" O(N^2) transformer attention is you correlate every token to every other token. So, in the limit, your network won't "miss" much.
When you abandon O(N^2) attention, you are forced to start adding heuristics to choose what to correlate. Any time you see one of those giant context window LLMs, you need to be asking what heuristics they added, what is getting correlated, and what is not getting correlated.
This paper chooses an exponential heuristic where tokens further in the past get exponentially less attention. This heuristic is fine for certain tasks like responding in a chat room, where the most recent tokens are the most important, but bad for tasks where tokens are roughly equally important throughout the text, such as a dense academic paper or a reference manual.
The bitter lesson [1] is going to eventually come for all of these. Eventually we'll figure out how to machine-learn the heuristic rather than hard code it. Recurrent neural networks (RNNs) do this implicitly, but we don't yet know how to effectively train RNNs on ultra-deep sequences.
Another possibility is learning a heuristic for non-recurrent LLMs via reinforcement learning, such as in [2], which is basically a reinforcement learned "auto-researcher" that was trained in a style reminiscent of AlphaGo.
[1] http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
[2] https://arxiv.org/pdf/2109.00527.pdf https://arxiv.org/pdf/2109.00527.pdf
- CuriouslyC 3y agoIt seems like building a context tree with a convex branch cross attention estimator then using branch and bound to prune the tree while descending to get exact cross attention when it's above a threshold would work pretty well, assuming the cross attention matrix actually is very sparse and the trouble is just accurately guessing the non-sparse elements.
- zzzzzzzza 3y agothis sounds to me like a dollar cost averaging strategy - only buy in when the current price falls below an n-day moving average. I doubt there is any risk adjusted alpha to the strategy - in practice it's my, newbie, understanding that the only thing that differentiates such strategies in the broader scheme of things is tax efficiency. however I am also not a ML expert
- linuxdude314 3y agoWhat are you talking about? Wrong thread?
- zzzzzzzza 3y agoi am suggesting the two strategies might have similar trade offs/benefits though I am not familiar enough with attention mechanisms to say for sure. it's a comparison/analogy?
- phillipcarter 3y agoThis comment makes so much sense relative to what I've seen with Claude's 1M context window. It reliably fails to succeed a task with a prompt where I just stuff in a big blob of data in the middle as context. But when I use emebddings to only select a small relevant subset of that data, it always passes the task.
- gamegoblin 3y agoYes, Claude 1M is using all sorts of approximation tricks to get that 1M context window. IMO this is actually quite deceptive marketing.
- dpflan 3y agoYes, that's the point now for competing AI research-for-profit companies, whatever metric is technical and sounds important, is going to be used in marketing and valuation determinations. It will be explored for research I'm sure, and then determine its product viability. It's nice competition, but agree, that it can be deceptive.
- l1n 3y agoClaude's context is 100K not 1M [1]. If you're somehow shoving in a million tokens that could explain the issue you're having! [1] https://www.anthropic.com/index/100k-context-windows https://www.anthropic.com/index/100k-context-windows
- gamegoblin 3y agoMisremembered, the main thrust of the comment still stands, the 100K context window isn't "real", it would be absurdly expensive to do it for real. They are using a lot of approximation tricks to get there.
- phillipcarter 3y agoYep, mistype on my end as well. Claude just fails to process the request if you get above 100k tokens (I've done that, heh).
- littlestymaar 3y agoNot disagreeing with your comment in general, but this particular sentence annoys me a bit: > where tokens are roughly equally important throughout the text, such as a dense academic paper or a reference manual. Even in these, not all tokens are equal, most of a text is actually pretty low-information, with key packs of token that contain most of the information that you're going to need throughout the entire text (that's why we use highlighters when learning). And that's why those O(n²) attention are pretty wasteful, at the same time, you need to be able to pick the proper token, and I agree with you that picking them through simple heuristic is probably not going to be enough.
- gamegoblin 3y agoBetter phrasing would have been "the important tokens are roughly evenly distributed throughout the text", that was the intended reading.
- bee_rider 3y agoAre you thinking more like a research paper or more like a textbook? For a textbook at least, it often seems to be the case that you need to have fully ingested the big picture ideas of one chapter to move on to some later ones, but this seems to me at least more like updating your model, rather than sampling context from the whole book (I mean it is an analogy of course, so neither matches perfectly).
- itissid 3y agoI would like to take a parallel view to the BitterLesson and how its playing out. There are exceptions. Its not only computation but also a mix of: 1. decades of theoretical breakthroughs coming together. 2. there is also collective human creativity and perseverance. Like Yann Le Cunn, Geoff Hinton etc have been working since the 90's and there were several milestones that were hit and it only caught on fire/went on steroids once the application(and the associated funding) was found due to creativity in the tech sector. But if the computation was somehow available before I am not sure it would have happened so quickly. Another example is that all methods under the AI umbrella are not dependent on crazy amounts of computation and data. Take the field of AutoRegressive models in Social/Life Sciences field. For example lets look at the STAN which broadly does heirarchical Bayesian Inference using MonteCarlo based methods in social science. It took some hard theoretical advancements to move the needle on MonteCarlo Simulation methods like detecting convergence and ability to have non conjugated priors for posterior sampling to work etc. The new methods are better by leaps and bounds over the conventional methods in the field. The computation for running the modern models from 2013 would be enough to run em for most cases.
- sashank_1509 3y agoBoth your points are not really valid. There have been decades of theoretical breakthroughs in computational linguistics too (Have there been any in Deep Learning?). There has also been a large amount of human creativity and perseverance in computational linguistics, arguably more than the amount I have seen in Deep Learning. Yet, not one useful algorithm has come from linguistics. In fact the old adage on speech processing can be applied to Natural Language Processing: "Every time I fire a linguist my performance improves by a few percent". The bitter lesson is bitter and important to keep in mind exactly because human creativity and perseverance do not matter in front of it. Consistently, the only methods that work are those that scale with computation, everything else does not matter. I would take a more extreme view, if computation didn't follow Moore's law, we wouldn't have invented alternate methods that do not require massive computation, we would just simply fail to do even the most basic tasks of intelligence and be stuck in the 1960s. A scary thought, but a true one I reckon. If computation kept following Moore's law but a few stalwarts like Yann Le Cun etc didn't exist, we would likely have found alternative architectures that scale and work, maybe not as good as ConvNets but transformers aren't as good as ConvNets either, they just need to scale.
- euclaise 3y ago> The bitter lesson [1] is going to eventually come for all of these. Eventually we'll figure out how to machine-learn the heuristic rather than hard code it. Recurrent neural networks (RNNs) do this implicitly, but we don't yet know how to effectively train RNNs on ultra-deep sequences. Linear RNNs and RWKV are examples of RNNs on deep sequences: https://arxiv.org/abs/2303.06349 https://arxiv.org/abs/2303.06349 https://arxiv.org/abs/2305.13048 https://arxiv.org/abs/2305.13048
- gamegoblin 3y agoI think the jury is still out if these will actually scale to ultra-long language understanding sequences. KWKV, for example, is still trained like GPT, but is architected so it can be run as an RNN during inference time. This is awesome, but it is unclear if the training regime will limit the effective use of long-ranging recurrent context.
- euclaise 3y agoTraining as GPT vs RNN will give you numerically identical results with RWKV, it's just two ways of computing the same thing. It's trained in GPT-mode because it's cheaper to train that way -- you can parallelize over the sequence length. In practice it isn't going to be any different than training with back-propagation through time for the same sequence length.
- sdenton4 3y agoThe work out of that group, starting with S4 layers, is 10000% the stuff to be paying attention to. https://srush.github.io/annotated-s4/ https://srush.github.io/annotated-s4/ HiPPO was brilliant - instead of working with the raw sequence, you work with its weighted laplace transform, and instead of actually computing the laplace transform you find the rule to update it when new data is added. Furthermore, we can 'band limit' the Laplace transform (similar to PCA) to keep only the 'most important' information while still preserving most of the information in the sequence - this is a common and quite effective compression technique. Any 'fast' transformer is going to be working with some kind of sampling or aggregation or compression of the long sequence. Sampling is ultimately going to be too noisy, and standard aggregations are going to be too coarse. So the thing to bet on is better compression techniques, which is what the S4/RWKV group are ultimately working on.
- mochomocha 3y agoWhile I agree with the beginning of your post, you lost me here: > The bitter lesson [1] is going to eventually come for all of these. Eventually we'll figure out how to machine-learn the heuristic rather than hard code it. Inefficiently re-learning over and over patterns that can be more explicitly encoded as smart inductive biases for better sample efficiency is what ML research is. The "bitter lesson" doesn't mean "throw the towel and your brain and just buy more GPUs". It means that the inductive biases / modeling strategies that win will always be the ones that are more hardware-friendly.
- gamegoblin 3y agoI agree with you that learning certain things is wasteful. For instance, one could imagine an RNN that learned to do some approximation of tree search for game playing Chess and Go. But we have very good reason to think that tree search is basically exactly what you want, so even systems like AlphaGo have the tree search implemented outside the neural net, but still using a learned system to heuristically guide the tree search. The reference to the bitter lesson here is that feature engineering has, thus far, typically lost out to more general end-to-end methods in the long run. This paper tries to do feature engineering by hand-coding an exponentially decaying mechanism, where tokens further in the past are assumed to be less important. My comment is that this type of hand-engineering will lose out to methods that are more end-to-end learned. These methods do not necessarily need to be hugely computationally intensive ("buy more GPUs"). That said, I could see it being the case that in the short term, we do just buy more GPUs, learn a general end-to-end algorithm, but eventually figure out how to re-implement that end-to-end learned algorithm in code significantly more efficiently.
- famouswaffles 3y agoBy and large, we don't really know what inductive biases we ought to be shoving in to models. Sometimes we think we do, but we're wrong more often than not. So methods with the least inductive biases work better.
- hotstickyballs 3y agoRecurrent networks are exponential. They also blow up and decay exponentially. So this is not necessarily worse than rnns
- AndrewKemendo 3y agoHaving studied Sutton for a long time now, what I take away from the bitter lesson is that the only pathway to Generally capable agents is to have the same scale of computational capacity in an embodied system as humans or other intelligent systems have. It’s effectively a product of physics and so we keep trying to outsmart physics is what Suttons point is - and you just can’t outsmart physics So, while the method probably is important in terms of efficiency or functionality within the current state of technological systems, the method is less important than the scale, and we’re not even close to the scale necessary yet.
- sdenton4 3y agoI don't think it's obvious that we don't have sufficient computational scale already... The human brain has ~86 billion neurons, but they only fire at like 2Hz, so you get 192 billion firings per second. GPT-3 has 175 billion parameters, and can apply those parameters much faster than the brain can fire neurons. Lots of folks like to point out that there's more complexity in neurons than model weights, which is fine, but it's not clear what the impact of that additional complexity actually is. Extra slow channels (eg, hormonal signals) also aren't an impossibility. So /maybe/ the models need to scale more, or maybe we need to figure out some better combination of training tasks to get the next big leap. There's massive progress being made with multi-modal inputs (which helps the model create a more coherent world-model, by relating text to images, audio, or video). Data selection. - picking good data instead of just throwing in everything - is also showing lots of promise. I tend to think there's some need for more 'interactive' or 'social' component to training, eg active/online learning with robotics - and figuring out how to get models to 'play'. Unstructured play is an essential mechanism in smart animals associated with hormonal signals and rewards - it's important, and we don't really know how to harness it yet. But overall, I don't think we're yet at a local maximum. There's too much cool stuff going on and the iron is still quite hot.
- AndrewKemendo 3y ago“Embodied” being one of the key things that you’re ignoring The brain =\= a human agent You need sensors and effectors and a highly variable motor system You can’t be “generally intelligent” if you do not have boundaries on your computing system which are mobile and have independent actions. In order to perform as well if not better than a human then you need to perform as well, if not better than a human in all possible environments, and those include the top of the world, the bottom of the ocean, every factory line, flying airplanes, etc…
- antonevstigneev 3y ago[dead]
- novok 3y agoOne clever trick I’ve seen is where before the text the goes out of the context window, it gets summarized by the llm and then that smaller summary is put into the context window and is continuously updated. It also reminds me how human memory works too.
- im3w1l 3y agoThese models seem to be able to cope with absolutely massive training sets, wheras the prompt input has to be quite small in comparison. I wonder if could leverage this state of affairs by shifting the prompt from input to training data. Like take a generic model, run a little bit of fine tuning on the prompt.
- DennisP 3y agoIt doesn't sound to me like it's quite "tokens further in the past get exponentially less attention." What they say is "attention allocation decreases exponentially as the distance between tokens grows." Instead of being quadratic because every pair of tokens gets the same attention, the tokens farther apart from each other get exponentially less. It doesn't matter how far they are from the final token. This seems to me more like a general computational approach than a hand-coded heuristic. David Shapiro claims it's similar to how the brain works, and has a neat analogy for it here: https://www.youtube.com/watch?v=R0wBMDoFkP0 https://www.youtube.com/watch?v=R0wBMDoFkP0
- refulgentis 3y agoThis is intriguing but I don't quite follow - really naive, but: isn't the final token as some position N? And given context size limit Y, when we generate the next token, right now I get attention from N - Y to N? And this supposes I get attention from 0 to N, but the attention decreases exponentially as we approach token 0?
- thomasahle 3y ago> Any time you see one of those giant context window LLMs, you need to be asking what heuristics they added, what is getting correlated, and what is not getting correlated. Exactly. The paper doesn't even contain any experiments with context windows over 32K tokens. Presumably because it doesn't really attend to the rest of those tokens at all. In practice it's just a 32K attention window with some "theoretical" opportunity for attending a bit further than that.
- smaddox 3y ago> Recurrent neural networks (RNNs) do this implicitly, but we don't yet know how to effectively train RNNs on ultra-deep sequences. What would you call "ultra-deep"? [1] shows how to train an RNN in a GPT-like mode, using parallel scan, with great performance on Path-X, which has a sequence length of 16k. It's based on prior papers doing the same thing but from a state space model perspective. [1] https://arxiv.org/abs/2303.06349 https://arxiv.org/abs/2303.06349
- quickthrower2 3y agoDo you get sparsity issues though? Say with a million tokens, what does attention between token 123887 “the” and token 4 “ing” even mean? You probably want less density of connections? Genuine question.
- eru 3y ago> When you abandon O(N^2) attention, you are forced to start adding heuristics to choose what to correlate. Any time you see one of those giant context window LLMs, you need to be asking what heuristics they added, what is getting correlated, and what is not getting correlated. Well, having a small context window and everything correlated with everything else is equivalent to having a large context window, but a particularly dumb heuristic.
- meghan_rain 3y agoGood point