4 ms·
I agree fine-tuning for task will give better results. Cerebras actually showed some research recently on this front. Sparse pre-training and dense fine-tuning
by boomerang90 4y ago
I agree fine-tuning for task will give better results. Cerebras actually showed some research recently on this front. Sparse pre-training and dense fine-tuning (https://arxiv.org/abs/2303.10464 https://arxiv.org/abs/2303.10464). You can recover the accuracy of sparse pre-trained models with dense fine-tuning and reduce FLOPs of the end-to-end pipeline by 2.5x compared to dense.