4 ms·How?by SknCode 7mo agoHow?sigmoid10 7mo agoSame way you distill any model. Training data efficiency matters only while you train the source model/ensemble. Once you have that you are purely compute bound during distillation.