3 ms·
I like the direction of the research of working in latent space but feeding the last layer representation back as a first layer embedding feels sketchy to me. T
by fabmilo 2y ago
I like the direction of the research of working in latent space but feeding the last layer representation back as a first layer embedding feels sketchy to me. Those layers have different representation space.
- zxexz 2y agoFeeding the last layer back as the input embedding has been done many times, e.g. Transformer-XL. The models are trained like this, it's not like they're taking a pre-trained Llama and just feeding it to itself. It's a simple, computationally cheap mechanism to add feedback.
- empath75 2y agoI read a paper not long ago that showed that deleting, duplicating and reordering layers doesn't actually seem to matter that much and it feeding back is just a kind of re-ordering.
- TeMPOraL 2y agoSo you're saying that feeding the last layer back to the first makes the model layer-order independent, or kinda infinitely deep, if you squint? :).
- torginus 2y agoImo this kind of makes sense - LLMs without a feedback loop can learn to have one themselves by encoding information in the previously generated tokens.
- imtringued 2y agoThey can't, because that would increase training loss. The training loss acts as a gatekeeper for reasoning.
- fabmilo 2y agofrom my understanding that is what they do, see the paper: > We use a pre-trained GPT-2 (Radford et al., 2019) as the base model for all experiments. I agree the feedback is necessary, and the mechanism simple and cheap, but I don't think is optimal.
- zxexz 2y agoYes, they use a pre-trained model, but they do further training (please correct me if I mis-read, and also I realize my above comment could be interpreted as saying they train a new model entirely from scratch). > We use a pre-trained GPT-2 (Radford et al., 2019) as the base model for all experiments. The learning rate is set to 1 × 10−4 while the effective batch size is 128. Following Deng et al. (2024), we also reset the optimizer when the training stages switch.
- liuliu 2y agoNot really. See the literature on sharing lm_head (last matrix multiplication) with the input embedding dict. Basically, the lm_head (a MxN matrix where M is the dictionary size and N is the internal dimension) can be seen as the dictionary too. You can think that and the softmax over it as compute cosine similarity of the last hidden output w.r.t. input embedding dictionary. In that sense, they are sharing the representation space. (BTW, I believe sharing lm_head with input embedding not working as good as separating them, so only mobile focused LLMs do so. So here is that. It would be interesting to experiment if injecting a projection layer like you suggested would improve performance or just red-herring).
- jsenn 2y ago> Those layers have different representation space. Do they? Interpretability techniques like the Logit Lens [1] wouldn't work if this were the case. That author found that at least for GPT-2, the network almost immediately transforms its hidden state into a "logitable" form: you can unproject the hidden state of any layer to see how that layer incrementally refines the next token prediction. [1]: https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreti...
- mbowcut2 2y agoThis was my first thought too. AFAIK each layer encodes different information, and it's not clear that the last layer would be able to communicate well with the first layer without substantial retraining. Like in a CNN for instance, if you fed later representations back in to the first kernels they wouldn't be able to find anything meaningful because it's not the image anymore, it's some latent representation of the image that the early kernels aren't trained on.
- danielmarkbruce 2y agollama 3.x is already sharing the last layer with the embedding layer, it just uses the transpose in the last layer operation.
- paraschopra 2y agoThe point is that training regime can force the network to immediately reshape the representation layer (after inputs) depending on whether it is a thought or language context.