3 ms·
The goal is distillation is to distill into smaller models like 7B, 1.5B. They didn't even change the model size, let alone try a different class of models. G
by rlforllms 2y ago
The goal is distillation is to distill into smaller models like 7B, 1.5B.
They didn't even change the model size, let alone try a different class of models.
Getting expert model's trajectories is trivial if you have vLLM to do batched inference.