10 ms·
Please look at these number with the following grain of salt, when optimizing programs. They are trumped by cache usage patterns [1]: Loading (fig IV)
by BenoitP 3y ago
Please look at these number with the following grain of salt, when optimizing programs. They are trumped by cache usage patterns [1]:
Loading (fig IV)
Location Energy (pJ = pico Joules)
L1 64 pJ/Byte
L2 121 pJ/Byte
L3 254 pJ/Byte
RAM 1250 pJ/Byte
Adding integers with SIMD (fig VI)
428 pJ/op, for 8 Byte/op; this means:
53 pJ/byte
So it takes actually 20 times to fetch data from RAM than to add it to something! And most often this is also the source of latency. Generally energy is linear to the distance signal has to travel, and that's the same for latency.
That's why successful data structures are sized a tiny bit under L1/L2 sizes! (BTree chunks, ring buffers).
If you've been following hardware, it's all about putting RAM closer to compute at the moment with chiplets at the moment.
[1] https://tu-dresden.de/zih/forschung/ressourcen/dateien/abgeschlossene-projekte/benchit/2010_IGCC_authors_version.pdf?lang=en https://tu-dresden.de/zih/forschung/ressourcen/dateien/abges...