3 ms·
In my experience, the most important aspect missing in most CPU GPU discussions, is that CPUs have a massive cache compared to GPUs, and that cache has pretty g
by marmaduke 6y ago
In my experience, the most important aspect missing in most CPU GPU discussions, is that CPUs have a massive cache compared to GPUs, and that cache has pretty good bandwidth (~30 GB/core?), even if main memory doesn't. So even if your task's hot data doesn't fit in L2 but in L3/core, AVX-whatever per core processing is a good bet regardless of what a GPU can do.
Another aspect that seems like a hidden assumption in CPU-GPU discussions is that you have the time-energy-expertise budget to (re)build your application to fit GPUs.
- dragontamer 6y agoOn the memory perspective, I basically see problems in roughly the following grouping of categories: 40TBs+ -- Storage-only solutions. "External Tape Merge sort algorithm", "Sequential Table Scan", etc. etc. (SSDs or even Hard drives if you go big enough) 4TB to 40TBs -- Multi-socket DDR4 RAM is king (8-way Ice Lake Xeon Scalable Platinum will probably reach 40TBs). Single-node distributed memory with NUMA / UPI to scale. 1TB to 4TB -- Single Socket DDR4 RAM (EPYC, even if at 4x NUMA. Or Single-node Ice Lake). 80GB to 1TB -- DGX / NVlink distributed memory A100 ganging up HBM2 together. GPU-distributed RAM is king. 256MBs to 80GBs -- HBM2 / GDDR6 Graphics RAM is king (80GB A100 2TB/s). 1.5MBs to 256MBs -- L3 cache is king (8x32MBs EPYC L3 cache, or POWER9 110MB+ L3 cache unified) 128kB to 1.5MBs -- L2 cache is king (1.25MB Ice Lake Xeons L2, this article) 1kB to 128kB -- L1 cache is king. (128kB L1 cache on Apple M1). Note: "GPU __Shared__" is a close analog to L1 and competes against it, but is shared between 32 to 256 GPU threads, so its not an apples-to-apples comparison. 1kB and below -- The realm of register-space solutions. (See 64-bit chess engine bitboards and the like). Almost fully CPU-constrained / GPU-constrained programming. 256x 32-bit GPU registers per GPU-thread / SIMD thread. CPUs have fewer nominal registers, but many "out of order" buffers or "reorder buffers" that practically count as register storage in a practical / pragmatic sense. CPUs just use their "real registers" as a mechanism to automatically discover parallelism in otherwise single-thread written code. ------------ As you can see: GPUs win in some categories, but CPUs win in others. And these numbers change every few months as a new CPU and/or GPU comes out. And at the lowest levels: CPUs and GPUs cannot be compared due to fundamental differences in architecture. For example: GPU __shared__ memory has gather/scatter capabilities (the NVidia PTX instructions / AMD GCN instructions permute vs bpermute), while CPUs traditionally only accelerate gather capabilities (pshufb), and leave vgather/vscatter instructions to the L1 cache instead. GPUs have 32x ports to __shared__, so every one of the 32-threads in a wave-front can read/write every single clock-tick (as long as all 32 they are on different ports/alignment, or you have a special one-to-all broadcast). CPUs only have 2 or 4 ports, so vscatter and vgather operate slowly, as if a single thread were reading/writing each of the memory locations. But CPU L1 cache has store-forwarding, MESI + cache coherence, and other acceleration features that GPUs don't have. GPUs are therefore more efficient at sharing data within workgroups of ~256 threads, but CPUs are more efficient at sharing data between cores, or even among out-of-die NUMA solutions, thanks to robust MESI messaging.