3 ms·
Nice! Can you explain what you mean by "simulate training beyond the number of available tokens"? Why does using distillation from a larger model simulate trai
by jakobov 2y ago
Nice! Can you explain what you mean by "simulate training beyond the number of available tokens"?
Why does using distillation from a larger model simulate training with more tokens?
- canyon289 2y agoHi, I work on the Gemma team (same as Alek opinions are my own). Essentially instead of tokens that are "already there" in text, the distillation allows us to simulate training data from a larger model
- suryabhupa 2y agoSurya here from the core Gemma team -- we can think of a distillation loss as learning to model the entire distribution of tokens that are likely to follow the prefix thus far, instead of only the token in the training example. If you do some back of the envelope calculations, we can see that learning to model a larger distribution yields many more bits of information to learn from.
- jakobov 2y agoGotcha. That makes sense. Thanks! What are the theories as to why this works better than training on a larger quantity of non-simulated tokens? Is it because the gradient from the non-simulated tokens is too noisy for a small model to model correctly?