2 ms·
"Pruning" is the main term you are looking for, with a variety of methods to do so. https://blog.dataiku.com/making-neural-networks-smaller-for-better-deployme
by snerbles 4y ago
"Pruning" is the main term you are looking for, with a variety of methods to do so.
https://blog.dataiku.com/making-neural-networks-smaller-for-better-deployment-solving-the-size-problem-of-cnns-using-network-pruning-with-keras https://blog.dataiku.com/making-neural-networks-smaller-for-...
https://arxiv.org/abs/2301.00774 https://arxiv.org/abs/2301.00774
https://tivadardanka.com/blog/how-to-compress-a-neural-network https://tivadardanka.com/blog/how-to-compress-a-neural-netwo...
There's also quantization. It's currently where a lot of the current grassroots research is happening, if you've seen the llama.cpp repo posted here on HN, it uses the original 16-bit float weights downsized to 4-bit integers with a comparable reduction in resource usage with relatively little performance loss.
The big players are also using quantization in hardware, most notably with Google's TPUs.