5 ms·
I’ve tried similar experiments before by asking the LLM to generate “internal” and “external” dialog, which I think is sort of the same idea at a higher level—-
by jacobsimon 2y ago
I’ve tried similar experiments before by asking the LLM to generate “internal” and “external” dialog, which I think is sort of the same idea at a higher level—-and might be preferable because it would allow for easy introspection vs a new set of tokens? I’m not enough of an expert to understand whether this proposal is intended more for training or inference.
- fesens 2y agoThe main advantage of using a new and constant token for reasoning is that, while we would pay the full price during training, in the inference phase, we could do most, if not all, the "reasoning" in one shot, without having to feed one generation token at a time.
- jacobsimon 2y agoCool!
- sdenton4 2y agoThis method is for training. They are using a stop-gradient to 'shield' some tokens from contributing to prediction of the immediate next token, and thus producing a stream of tokens that are only used for longer term prediction. This is a bit more low level than the usual prompt engineering approaches, and to my mind, a bit more promising. There's more easily measurable results, and I've seen other context where a well placed stop-gradient does wonders...
- lucidrains 2y agoyes, it is a stop gradient mask on the attention matrix, iiuc. worth trying
- lucidrains 2y agocould even try it with a fraction of the attention heads, instead of introducing new tokens
- sdenton4 2y agoAn important piece here is that there's still a training signal making it to the makes weights. See SimSiam for a similar example.
- lucidrains 2y agoindeed, simsiam is a great example of the effectiveness of using stop gradient