4 ms·
Reminds me of another Tencent paper https://dl.acm.org/doi/10.1145/3711896.3736949 https://dl.acm.org/doi/10.1145/3711896.3736949 that is how to combine distill
by willvarfar 1y ago
Reminds me of another Tencent paper https://dl.acm.org/doi/10.1145/3711896.3736949 https://dl.acm.org/doi/10.1145/3711896.3736949 that is how to combine distillation and ensemble for faster parallel inference.
That was Tencent doing parallelism at the model level. And now this is their evolution on MoE. Very complementary.