3 ms·
"...because it has to somehow tell the next layer how to generate the next token prediction." -- This isn't actually true in the case of transformers. Features
by psb217 2y ago
"...because it has to somehow tell the next layer how to generate the next token prediction." -- This isn't actually true in the case of transformers. Features in the final TF layer at time t in a sequence do not depend on the features in the final TF layer at any other time step. Recurrence in transformers is done "depthwise" via "causally masked" convolutions. Final layer features at time t can depend on penultimate layer features at time t-1, but not on final layer features at time t-1.
- danielmarkbruce 2y agoyou are misunderstanding what the person is saying. They are saying the final hidden layer outputs a vector which has all the information that decides the logits which decide the probabilities of each token in the entire vocabulary. Ie, it is storing a lot of information.
- ttul 2y agoCorrect. And although the final layer outputs a softmax of the token probabilities, the model by that point surely has a rich understanding of more than just the next token it wants to predict.
- danielmarkbruce 2y agoYup, some tokens are effectively branching decisions. Yann has a whole rant about a shortcoming of LLMs being they take the same compute regardless of the position in a sentence - which isn't great because sometimes you really have a serious decision to make, other times not so much. It also makes you wonder about optimal embedding size - maybe the right size is 10x bigger.
- versteegen 2y ago> surely has a rich understanding of more than just the next token it wants to predict > the last hidden layer is obviously super rich in semantic information I don't agree that this is obvious, and think it's likely wrong (see the sibling thread [1]). The model has to at some point compress down its prediction for the entire future string of text to a prediction for a single token. There's no prior reason to assume it does this mostly in the final "LM head" linear layer, and the inputs to it don't have to predict anything other than the very next token so there's no reason it should (which is what I think psb217 was getting at), but I'm not familiar with what research has been done into it. On the other hand, processing seems to typically be concentrated in the central layers. [1] https://news.ycombinator.com/item?id=42379167 https://news.ycombinator.com/item?id=42379167
- danielmarkbruce 2y agoThe last hidden layer outputs a vector which is then used to predict the probabilities of every token in the vocabulary, by a single layer (and, in practice now in llama models, this layer is the transpose of the embedding layer). That vector has a lot of information in it, it's not a debatable thing. As noted above in parens, look at the llama 3.x models. The space is already shared in some sense. It's called "tied embedding".
- versteegen 2y ago> That vector has a lot of information in it, it's not a debatable thing. Encoding the next token is the minimum possible amount of information it might contain; that's not much information (the distribution over the next token is just a projection from the embedding space). E.g. it would be useless for any classification task.
- danielmarkbruce 2y agoVarious models in production are doing exactly that - training a layer which takes the vector out of the last hidden layer, for classification, in place of the language head. I even have one in production right now doing regression using the output of the last hidden layer.... In the case of llama 3 its 4096 * 16 bit = 8192 bytes of information...that's like 8192 characters of ascii. More than enough for most classification tasks... and if you jsut spend any time thinking about encoding the logits for a vocab of 128k... you'll come to the conclusion it's likely to require at least several hundred bytes (maybe 1000?) to do it in any way that will actually work in practice.
- ttul 2y agoThink of it like this: the final softmax layer is like being forced to pick a single word as your next prediction, while the hidden layer contains all the reasoning and understanding that led to that decision. It's similar to how a human might have a complex thought but needs to reduce it to a single word when speaking. Many existing applications make use of hidden layers in a transformer to perform useful tasks such as classification. The concept of an “embedding” is simply the output of a hidden layer, after all.