3 ms·
I think the key thing is that you assume a single dot-product per RAM load. In general you get higher compute performance by doing more than this. You load weig
by Joky 7y ago
I think the key thing is that you assume a single dot-product per RAM load. In general you get higher compute performance by doing more than this. You load weights for a layer in the SRAM, then stream from RAM tiles of data that you can compute multiple operations on.
In the same way, to reach peak FLOPs on a GPU you better use the local/shared memory as much as possible.