7 ms·
Scaling Transformers to 1B Tokens
- bratao 3y agoI need to carefully read the article, but sparse attention is an interesting technique that has been used previously (as in BigBird) but has often proved to perform (way) worse than full attention. The sliding component that performs full attention is indeed useful (much like the Blockwise Parallel Transformer), but the sparse patterns are elements that don't intuitively resonate with me. The model might select random words in the context. There's definitely a case where this could be unfortunate if it ends up selecting irrelevant words. The graph on the first page, in my opinion, seems like a needless flex
- londons_explore 3y ago> The graph on the first page, in my opinion, seems like a needless flex Indeed - they used half of the cover page of their paper to show a chart which illustrates... nothing...
- gamegoblin 3y agoThe benefit of "traditional" O(N^2) transformer attention is you correlate every token to every other token. So, in the limit, your network won't "miss" much. When you abandon O(N^2) attention, you are forced to start adding heuristics to choose what to correlate. Any time you see one of those giant context window LLMs, you need to be asking what heuristics they added, what is getting correlated, and what is not getting correlated. This paper chooses an exponential heuristic where tokens further in the past get exponentially less attention. This heuristic is fine for certain tasks like responding in a chat room, where the most recent tokens are the most important, but bad for tasks where tokens are roughly equally important throughout the text, such as a dense academic paper or a reference manual. The bitter lesson [1] is going to eventually come for all of these. Eventually we'll figure out how to machine-learn the heuristic rather than hard code it. Recurrent neural networks (RNNs) do this implicitly, but we don't yet know how to effectively train RNNs on ultra-deep sequences. Another possibility is learning a heuristic for non-recurrent LLMs via reinforcement learning, such as in [2], which is basically a reinforcement learned "auto-researcher" that was trained in a style reminiscent of AlphaGo. [1] http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html [2] https://arxiv.org/pdf/2109.00527.pdf https://arxiv.org/pdf/2109.00527.pdf
- CuriouslyC 3y agoIt seems like building a context tree with a convex branch cross attention estimator then using branch and bound to prune the tree while descending to get exact cross attention when it's above a threshold would work pretty well, assuming the cross attention matrix actually is very sparse and the trouble is just accurately guessing the non-sparse elements.
- zzzzzzzza 3y agothis sounds to me like a dollar cost averaging strategy - only buy in when the current price falls below an n-day moving average. I doubt there is any risk adjusted alpha to the strategy - in practice it's my, newbie, understanding that the only thing that differentiates such strategies in the broader scheme of things is tax efficiency. however I am also not a ML expert
- linuxdude314 3y agoWhat are you talking about? Wrong thread?
- zzzzzzzza 3y agoi am suggesting the two strategies might have similar trade offs/benefits though I am not familiar enough with attention mechanisms to say for sure. it's a comparison/analogy?
- phillipcarter 3y agoThis comment makes so much sense relative to what I've seen with Claude's 1M context window. It reliably fails to succeed a task with a prompt where I just stuff in a big blob of data in the middle as context. But when I use emebddings to only select a small relevant subset of that data, it always passes the task.
- gamegoblin 3y ago
- cs702 3y agoWell, this looks promising. The key idea is to collect a different set of tokens, with different levels of sparsity for each head, apply regular (dense) self-attention over all heads, weighted by pairwise distance, and spread and add the output residuals to their corresponding location in the original sequence. It seems to work really well, judging by the perplexity scores shown in the paper -- though we don't yet know if those perplexity scores will translate into good performance on real-world tasks. I'm going to take a closer look.
- londons_explore 3y agoThey use perplexity on github data to demonstrate the effectiveness of their model. I suspect github data has a lot of copy pasted code. Ie. a good chunk of what you are asking the model to do is to go back X million tokens and copy a chunk verbatim. Sure, the model might also be looking back at some code X million tokens ago and using that to improve its guess of the next token (oh look, the API definition of the API I am using is back here, that'll help me get this right!). But the perplexity number alone doesn't differentiate those cases - and considering how much code copying/templating happens in software, I suspect that affects the perplexity a lot more than smartly using stuff from the context window. I wonder if these models work well on other kinds of data?
- deleted 3y ago[deleted]
- spuz 3y agoWhat does the "number of tokens" characteristic of an LLM mean exactly? How does 1B compare with GPT-3.5 or GPT-4?
- rising-sky 3y agoContext length / window. Think of them and the "number of words" that the model can effectively process. 1 token is roughly equal to 4 characters or 0.75 words for English text. The number of tokens is the total number that can fit into a context window, which again is the space of "input" i.e. prompts and output (response/ completions) that the model can handle
- PartiallyTyped 3y agoA sequence of characters is encoded into tokens, tokens are grouped characters, each token is mapped to a vector representation. When you give text to an LLM, the text is encoded into tokens, and each token corresponds to an index. Each index corresponds to one vector. The model produces vectors, and then finds the most similar vector and selects the corresponding index as the next token. This is a spectrum, you can write a model that works on the bit level, so 2 vectors, or byte level, 256, or pairs of bytes, 2^16 and so on and so forth. These days, we use statistical approaches to build the tokens, and a token can be 1, 2 or 3 or N characters long. So when you give a sequence of characters to the model, it turns that to a sequence of tokens and loads a vector for each one, and when doing computations, it needs to consider all tokens together. This is called the context window. In this case, scaling the number of tokens means scaling the context window to a large number. GPT3.5 can do 2Ki tokens iirc, OpenAI’s GPT4 can do 4Ki iirc, Claude from anthropic can do 1Mi iirc. The context window is kinda analogous to your working memory, the higher the better, unless there are approximations that trade off quality for length, which is what is happening here.
- treprinum 3y agoOriginal GPT3.5 can do 4k tokens and there is a recent version with 16k tokens (gpt-3.5-turbo-16k)
- Imnimo 3y agoWithout any experiment showing that language modeling performance actually continues to improve past 32k tokens using this scheme, how are we supposed to tell whether this is actually viable?
- euclaise 3y agoImportant note: They only did experiments up to 32k length
- WanderPanda 3y agoHow stupendous to not put the first figure on a log scale...
- oofbey 3y agoCute trick but the opposite of helpful. Is the goal of your paper to brag or educate?
- jeron 3y agothis paper leans towards the former
- quickthrower2 3y agoIt is almost an XKCD style piss-take.
- daemonk 3y agoThis is relevant also: https://hazyresearch.stanford.edu/blog/2023-03-07-hyena https://hazyresearch.stanford.edu/blog/2023-03-07-hyena
- kytazo 3y agoIs assuming the sequence length is directly correlated to the context window a meaningful thought? Does this imply similar increases in context in practice?
- jumpCastle 3y agoTitle with a 10 digits number, meaningless first page figure and no experiments related to the main claim. Did a rogue author posted it without permission again?
- jeron 3y agoAs noted, they only did experiments up to 32k length which is silly considering the title
- jumpCastle 3y agoSilly is a charitable interpretation.
- sillysaurusx 3y agoYes, it’s bunk. https://twitter.com/theshawwn/status/1676822953210662913?s=46&t=CHNZIj4cQ8XNT6qqZQipWQ https://twitter.com/theshawwn/status/1676822953210662913?s=4... (A researcher at Brain confirmed it’s not worth reading: https://twitter.com/giffmana/status/1676864336764055552?s=46&t=CHNZIj4cQ8XNT6qqZQipWQ https://twitter.com/giffmana/status/1676864336764055552?s=46...) I hate being dismissive, but I’ve been dragged to the conclusion that headlines matter in research, and people chase headlines. Three orders of magnitude jump, instantly, isn’t plausible. It almost never happens.
- climatologist 3y agoDoes anyone know if 1B tokens is enough to solve sudoku puzzles?