2 ms·
There seems to be some relevant prior work that is not referenced by this paper, such as our work on training sparse neural networks: https://arxiv.org/abs/171
by dpkingma 2y ago
There seems to be some relevant prior work that is not referenced by this paper, such as our work on training sparse neural networks:
https://arxiv.org/abs/1712.01312 https://arxiv.org/abs/1712.01312
Abstract: "We propose a practical method for L0 norm regularization for neural networks: pruning the network during training by encouraging weights to become exactly zero. Such regularization is interesting since (1) it can greatly speed up training and inference, and (2) it can improve generalization. [...]"