4 ms·
Compressing LLMs with progressive pruning and multi-objective distillation
- adam_patarino 6mo agoCompressing a mixture of experts model to fit on smaller hardware with a reinforcement learning approach called Self-Distillation Policy Optimization, progressive expert pruning, multi-objective knowledge distillation, speculative decoding, and custom quantization.
- MikeSynnott 6mo ago[dead]
- MikeSynnott 6mo agoAwesome!