5 ms·
I'm not saying this to be rude, but I think you have a deep misunderstanding of how AI training works. You cannot just skip the matrix multiplications necessary
by Invictus0 2y ago
I'm not saying this to be rude, but I think you have a deep misunderstanding of how AI training works. You cannot just skip the matrix multiplications necessary to train the model, or get current hardware to do it faster.
- xdavidliu 2y agowas the first sentence really necessary? The second sentence seems fine by itself.
- nickpsecurity 2y agoThere's work on replacing multiplication. Here's four examples: https://openaccess.thecvf.com/content_CVPR_2020/papers/Chen_AdderNet_Do_We_Really_Need_Multiplications_in_Deep_Learning_CVPR_2020_paper.pdf https://openaccess.thecvf.com/content_CVPR_2020/papers/Chen_... https://arxiv.org/abs/2012.03458 https://arxiv.org/abs/2012.03458 https://openaccess.thecvf.com/content/CVPR2021W/MAI/papers/Elhoushi_DeepShift_Towards_Multiplication-Less_Neural_Networks_CVPRW_2021_paper.pdf https://openaccess.thecvf.com/content/CVPR2021W/MAI/papers/E... https://arxiv.org/pdf/2106.10860 https://arxiv.org/pdf/2106.10860
- benterix 2y agoNo offence taken! As far as my (shallow!) understanding goes, the main challenge is the need for many GPUs with huge amounts of memory, and it still takes ages to train the model. So regarding the use of consumer GPUs, some work has been done already, and I've seen some setups where people combine of these and are successful. As for the the other aspects, maybe at some point we distill what is really needed to a smaller but excellent dataset that would give similar results in the final models.