3 ms·
Actually it's super easy, barely an inconvenience, since the raw memory bandwidth you get per core is around 1 byte per cycle and half that per thread. A 3900X
by blattimwind 7y ago
Actually it's super easy, barely an inconvenience, since the raw memory bandwidth you get per core is around 1 byte per cycle and half that per thread. A 3900X with 12 cores / 24 threads getting a total memory bandwidth of 57.6 GB/s (using DDR4-3600) has just 0.6 bytes / cycle and thread. Meanwhile the CPU can actually achieve ALU throughputs that are 20 times greater than that. Even encryption algorithms are much faster than that.
The saving grace is not that there is a lot of memory bandwidth to go around; there isn't; the saving grace is that a lot of processing either doesn't use too much data (making caches effective) or is complex enough to not be limited by memory bandwidth. Rarely using all cores and threads at the same time helps a lot as well.
- dragontamer 7y ago> Actually it's super easy, barely an inconvenience Wow wow wow. Your references are TIGHT. ---------- The main RAM / CPU issue is latency: it takes over 100-clock cycles to communicate with DDR4, which a CPU could process ~400 instructions if they were lined up just right (modern CPUs are super-scalar, executing multiple instructions per clock tick). Only once you have sizable caches + data locality will you solve the latency problem. Without locality, I'm not sure if the latency problem can be solved at all. Fortunately, many problems have an element of locality that can be taken advantage of.
- fulafel 7y agoMemory controllers do prefetching for streaming memory accesses. Yes a lot of workloads are more latency sensitive, but it's not like sequential streaming memory accesses are hard given the prefetching. (Anyone have pointers to current EPYC vs Xeon STREAM results?)
- BubRoss 7y agoThat might be what the parent is saying. Most software is so poorly written that even if it needs to do something CPU intensive it is probably still skipping around in memory so much that memory bandwidth isn't an issue. My experience is that it takes well structured memory access to actually max out memory bandwidth even if the numbers seem like the CPU would be starved.