3 ms·
> That vector has a lot of information in it, it's not a debatable thing. Encoding the next token is the minimum possible amount of information it might contai
by versteegen 2y ago
> That vector has a lot of information in it, it's not a debatable thing.
Encoding the next token is the minimum possible amount of information it might contain; that's not much information (the distribution over the next token is just a projection from the embedding space). E.g. it would be useless for any classification task.
- danielmarkbruce 2y agoVarious models in production are doing exactly that - training a layer which takes the vector out of the last hidden layer, for classification, in place of the language head. I even have one in production right now doing regression using the output of the last hidden layer.... In the case of llama 3 its 4096 * 16 bit = 8192 bytes of information...that's like 8192 characters of ascii. More than enough for most classification tasks... and if you jsut spend any time thinking about encoding the logits for a vocab of 128k... you'll come to the conclusion it's likely to require at least several hundred bytes (maybe 1000?) to do it in any way that will actually work in practice.