3 ms·
That is correct, if I understand you correctly, but that doesn't solve the entire optimization problem. You still have to figure out how exactly to handle tiles
by lacker 4y ago
That is correct, if I understand you correctly, but that doesn't solve the entire optimization problem. You still have to figure out how exactly to handle tiles and transfers to shared memory. This might be a good page for answering this question:
https://docs.nvidia.com/deeplearning/performance/dl-performance-matrix-multiplication/index.html https://docs.nvidia.com/deeplearning/performance/dl-performa...