7 ms·
Thanks for the feedback! Let me try to state it better: In the end, we only use the next-token head for generating. So which parts of the 2-token target H(X) +
by faabian 2y ago
Thanks for the feedback! Let me try to state it better:
In the end, we only use the next-token head for generating. So which parts of the 2-token target H(X) + H(Y) are "auxiliary" in the sense that they help learning and which are "wasted"? H(X | Y) and I(X; Y) are useful for next-token generation while, by definition, H(Y | X) is the information quantity not related to the next token X. So we could say: "multi-token prediction trades the useful information I(X; Y) from H(Y) for the wasted computations on H(Y | X)".
However, note that H(Y | X) is a next-token entropy for predicting Y from the prefix (C, X). If the attention mechanism allows to transfer computations already made for predicting Y|X to the next step, these computations may actually not have been wasted -- it was just pre-computations.
- stealthcat 2y agoDid you have some small toy experiments to prove this?