10 ms·
Do the resnet results use the new sparsity feature ? I'd be interested to know what impact does that have.
by fluffything 6y ago
Do the resnet results use the new sparsity feature ? I'd be interested to know what impact does that have.
- whatever1 6y agoWhat is this feature and can it work in generic matrices? It could be a game changer in physics and operations research where the matrices are sparse.
- hydroreadsstuff 6y agoHalf the entries in your matrices need to be 0, then the Hardware will compress them and execute the matrix-matrix multiplication 2x as fast. On a tensor instruction level.
- the_svd_doctor 6y agoI don’t think “sparsity” in ML, where like 10...50% is sparse, is the same as sparsity in Physics, where 99.999% (add more 9 with larger problems) of your matrix is sparse.
- ryneandal 6y agoIt should, AFAIK that was on the SM-level. > Ampere's benefit is that it can deal with dense and sparse matrices differently. Its cores are twice as fast as Turing's for dense matrix and four times as quick for sparse matrix that have all the needless weights removed. The upshot, per SM, is dense processing at the same speed - it has half the cores, remember - and twice the overall throughput for sparse processing. https://hexus.net/tech/reviews/graphics/145342-nvidia-geforce-rtx-3080-founders-edition-ampere/?page=2 https://hexus.net/tech/reviews/graphics/145342-nvidia-geforc...
- std_badalloc 6y agoYeah, but ResNet does not have sparse matrices, so how could it use them? Post ReLU activations may be sparse, but I don't think that helps when used with a non-sparse Conv2d.
- ryneandal 6y agoAh, I should have paid more attention to the question. Read it as "are they enabled?" My bad. :(
- dplavery92 6y agoI don't know if there are any white papers with hard details yet (if anyone knows of one, please share!), but nVidia's marketing material[0] for the Ampere architecture claims the following: "Sparsity is possible in deep learning because the importance of individual weights evolves during the learning process, and by the end of network training, only a subset of weights have acquired a meaningful purpose in determining the learned output. The remaining weights are no longer needed. Fine grained structured sparsity imposes a constraint on the allowed sparsity pattern, making it more efficient for hardware to do the necessary alignment of input operands. Because deep learning networks are able to adapt weights during the training process based on training feedback, NVIDIA engineers have found in general that the structure constraint does not impact the accuracy of the trained network for inferencing. This enables inferencing acceleration with sparsity." So the idea seems to be that at the end of training, there's fine tuning that can be done to figure out which weights can be zeroed out without significantly impacting prediction accuracy, and then you can accelerate inferences with sparse matrix multiplication. They consider training acceleration with sparse matrices an "active research area." I could see it being nice for the sake of running large language models on consumer, or really cool for the few edge computing applications that can actually demand and power conventional GPUs (e.g. self-driving cars.) It's probably not a great boon to the researcher who wants to reduce their iteration timeline though. [0] https://developer.nvidia.com/blog/nvidia-ampere-architecture-in-depth/ https://developer.nvidia.com/blog/nvidia-ampere-architecture...