8 ms·
> This is powerful evidence that even though models are trained to output one word at a time I find this oversimplification of LLMs to be frequently poisonous
by fpgaminer 2y ago
> This is powerful evidence that even though models are trained to output one word at a time
I find this oversimplification of LLMs to be frequently poisonous to discussions surrounding them. No user facing LLM today is trained on next token prediction.
- rco8786 2y agoSuper interesting. Can you explain more, or provide some reading? I’m obviously behind
- JKCalhoun 2y agoAs a layman though, I often see this description for how it is LLMs work.
- fpgaminer 2y agoRight, but it leads to too many false conclusions by lay people. User facing LLMs are only trained on next token prediction during initial stages of their training. They have to go through Reinforcement Learning before they become useful to users, and RL training occurs on complete responses, not just token-by-token. That leads to conclusions elucidated by the very article, that LLMs couldn't possibly plan ahead because they are only trained to predict next tokens. When the opposite conclusion would be more common if it was better understood that they go through RL.
- mentalgear 2y agoWhat? The "article" is from anthropic, so I think they would know what they write about. Also, RL is an additional training process that does not negate that GPT / transformers are left-right autoencoders that are effectively next token predictors. [Why Can't AI Make Its Own Discoveries? — With Yann LeCun] (https://www.youtube.com/watch?v=qvNCVYkHKfg https://www.youtube.com/watch?v=qvNCVYkHKfg)
- pipes 2y agoListening to this today, so far really good. Glad I found it. Thanks.
- TeMPOraL 2y agoYou don't need RL for the conclusion "trained to predict next token => only things one token ahead" to be wrong. After all, the LLM is predicting that next token from something - a context, that's many tokens long. Human text isn't arbitrary and random, there are statistical patterns in our speech, writing, thinking, that span words, sentences, paragraphs - and even for next token prediction, predicting correctly means learning those same patterns. It's not hard to imagine the model generating token N is already thinking about tokens N+1 thru N+100, by virtue of statistical patterns of preceding hundred tokens changing with each subsequent token choice.
- fpgaminer 2y agoTrue. See one of Anthropic's researcher's comment for a great example of that. It's likely that "planning" inherently exists in the raw LLM and RL is just bringing it to the forefront. I just think it's helpful to understand that all of these models people are interacting with were trained with the _explicit_ goal of maximizing the probabilities of responses _as a whole_, not just maximizing probabilities of individual tokens.
- losvedir 2y agoThat's news to me, and I thought I had a good layman's understanding of it. How does it work then?
- fpgaminer 2y agoAll user facing LLMs go through Reinforcement Learning. Contrary to popular belief, RL's _primary_ purpose isn't to "align" them to make them "safe." It's to make them actually usable. LLMs that haven't gone through RL are useless to users. They are very unreliable, and will frequently go off the rails spewing garbage, going into repetition loops, etc. RL learning involves training the models on entire responses, not token-by-token loss (1). This makes them orders of magnitude more reliable (2). It forces them to consider what they're going to write. The obvious conclusion is that they plan (3). Hence why the myth that LLMs are strictly next token prediction machines is so unhelpful and poisonous to discuss. The models still _generate_ response token-by-token, but they pick tokens _not_ based on tokens that maximize probabilities at each token. Rather they learn to pick tokens that maximize probabilities of the _entire response_. (1) Slight nuance: All RL schemes for LLMs have to break the reward down into token-by-token losses. But those losses are based on a "whole response reward" or some combination of rewards. (2) Raw LLMs go haywire roughly 1 in 10 times, varying depending on context. Some tasks make them go haywire almost every time, other tasks are more reliable. RL'd LLMs are reliable on the order of 1 in 10000 errors or better. (3) It's _possible_ that they don't learn to plan through this scheme. There are alternative solutions that don't involve planning ahead. So Anthropic's research here is very important and useful. P.S. I should point out that many researchers get this wrong too, or at least haven't fully internalized it. The lack of truly understanding the purpose of RL is why models like Qwen, Deepseek, Mistral, etc are all so unreliable and unusable by real companies compared to OpenAI, Google, and Anthropic's models. This understanding that even the most basic RL takes LLMs from useless to useful then leads to the obvious conclusion: what if we used more complicated RL? And guess what, more complicated RL led to reasoning models. Hmm, I wonder what the next step is?
- scudsworth 2y agofirst footnote: ok ok they're trained token by token, BUT
- SkyBelow 2y agoIgnoring for a moment their training, how do they function? They do seem to output a limited selection of text at a time (be it a single token or some larger group). Maybe it is the wording of "trained to" verses "trained on", but I would like to know more why "trained to" is an incorrect statement when it seems that is how they function when one engages them.
- sdwr 2y agoIn the article, it describes an internal state of the model that is preserved between lines ("rabbit"), and how the model combines parallel calculations to arrive at a single answer (the math problem) People output one token (word) at a time when talking. Does that mean people can only think one word in advance?
- sroussey 2y agoSome people don’t even do that!
- wuliwong 2y agoBad analogy, an LLM can output a block of text all at once and it wouldn't impact the user's ability to understand it. If people spoke all the words in a sentence at the same time, it would not be decipherable. Even writing doesn't yield a good analogy, a human writing physically has to write one letter at a time. An LLM does not have that limitation.
- sdwr 2y agoThe point I'm trying to make is that "each word following the last" is a limitation of the medium, not the speaker. Language expects/requires words in order. Both people and LLMs produce that. If you want to get into the nitty-gritty, people are perfectly capable of doing multiple things simultaneously as well, using: - interrupts to handle task-switching (simulated multitasking) - independent subconscious actions (real multitasking) - superpositions of multiple goals (??)
- SkyBelow 2y ago
- drcode 2y agoThat's seems silly, it's not poisonous to talk about next token prediction if 90% of the training compute is still spent on training via next token prediction (as far as I am aware)
- fpgaminer 2y ago99% of evolution was spent on single cell organisms. Intelligence only took 0.1% of evolution's training compute.
- drcode 2y agook that's a fair point
- diab0lic 2y agoI don’t really think that it is. Evolution is a random search, training a neural network is done with a gradient. The former is dependent on rare (and unexpected) events occurring, the latter is expected to converge in proportion to the volume of compute.
- jpadkins 2y agowhy do you think evolution is a random search? I thought evolutionary pressures, and the mechanisms like epigenetics make it something different than a random search.
- devmor 2y agoEvolution also has no "goal" other than fitness for reproduction. Training a neural network is done intentionally with an expected end result.
- rcxdude 2y agoThere's still a loss function, it's just an implicit, natural one, instead of artificially imposed (at least, until humans started doing selective breeding). The comparison isn't nonsense, but it's also not obvious that it's tremendously helpful (what parts and features of an LLM are analagous to what evolution figured out with single-celled organisms compares to multicellular life? I don't know if there's actually a correspondance there)
- pmontra 2y agoAnd no users which are facing a LLM today have been trained on next token prediction when they were babies. I believe that LLMs and us are thinking in two very different ways, like airplanes, birds, insects and quad-drones fly in very different ways and can perform different tasks. Maybe no bird looking at a plane would say that it is flying properly. Instead it could be only a rude approximation, useful only to those weird bipeds an scary for everyone else. By the way, I read your final sentence with the meaning of my first one and only after a while I realized the intended meaning. This is interesting on its own. Natural languages.
- naasking 2y ago> And no users which are facing a LLM today have been trained on next token prediction when they were babies. That's conjecture actually, see predictive coding. Note that "tokens" don't have to be language tokens.
- colah3 2y agoHi! I lead interpretability research at Anthropic. I also used to do a lot of basic ML pedagogy (https://colah.github.io/ https://colah.github.io/). I think this post and its children have some important questions about modern deep learning and how it relates to our present research, and wanted to take the opportunity to try and clarify a few things. When people talk about models "just predicting the next word", this is a popularization of the fact that modern LLMs are "autoregressive" models. This actually has two components: an architectural component (the model generates words one at a time), and a loss component (it maximizes probability). As the parent says, modern LLMs are finetuned with a different loss function after pretraining. This means that in some strict sense they're no longer autoregressive models – but they do still generate text one word at a time. I think this really is the heart of the "just predicting the next word" critique. This brings us to a debate which goes back many, many years: what does it mean to predict the next word? Many researchers, including myself, have believed that if you want to predict the next word really well, you need to do a lot more. (And with this paper, we're able to see this mechanistically!) Here's an example, which we didn't put in the paper: How does Claude answer "What do you call someone who studies the stars?" with "An astronomer"? In order to predict "An" instead of “A”, you need to know that you're going to say something that starts with a vowel next. So you're incentivized to figure out one word ahead, and indeed, Claude realizes it's going to say astronomer and works backwards. This is a kind of very, very small scale planning – but you can see how even just a pure autoregressive model is incentivized to do it.
- deleted 2y ago[deleted]
- stonemetal12 2y ago> In order to predict "An" instead of “A”, you need to know that you're going to say something that starts with a vowel next. So you're incentivized to figure out one word ahead, and indeed, Claude realizes it's going to say astronomer and works backwards. Is there evidence of working backwards? From a next token point of view, predicting the token after "An" is going to heavily favor a vowel. Similarly predicting the token after "A" is going to heavily favor not a vowel.
- boodleboodle 2y agoThis is why, whenever I can, I call RLHF/DPO "sequence level calibration" instead of "alignment tuning". Some precursors to RLHF: https://arxiv.org/abs/2210.00045 https://arxiv.org/abs/2210.00045 https://arxiv.org/abs/2203.16804 https://arxiv.org/abs/2203.16804