4 ms·
Literally retraining on inputs and outputs without doing anything clever, does not to me seem to deserve quite such a fancy name. How do we end up with a smalle
by Eliezer 4y ago
Literally retraining on inputs and outputs without doing anything clever, does not to me seem to deserve quite such a fancy name. How do we end up with a smaller model if the dataset formats are exactly similar, except in some other special case, like using much more generated data to train the smaller model? Of course there are other deservedly clever things you could do besides matching logits, that would enable a smaller student model.
Here, the Stanford authors are not doing anything clever to enable a smaller model. They're just yoinking the fine-tuning onto what happens to be a smaller model. The destination model being smaller is not the point. The cheap yoinking of just the instruction tuning is the point. They used a small model as the destination because that was cheaper.
- Imnimo 4y agoSure, but the value of logit matching is that it allows the smaller model to access richer representations that it has sufficient capacity to express but insufficient capacity to learn. In this case, we're not trying to distill the full capabilities of the large model. We just want to mimic instruction following, and we don't have any reason to believe the smaller model has insufficient capacity. So doesn't the lesson of knowledge distillation tell us that this should work just by matching low temperature logits (in other words, matching argmax only)? I don't see how this is fundamentally different from the standard knowledge distillation setting - it's just an easier instance of it