4 ms·
You're arguing at a particular granularity and even that doesn't hold uniformly. A64FX is HBM2 at equivalent bandwidth to GPUs (with lower power). CCX is much
by jedbrown 6y ago
You're arguing at a particular granularity and even that doesn't hold uniformly.
A64FX is HBM2 at equivalent bandwidth to GPUs (with lower power). CCX is much finer granularity than an entire GPU, so not a direct comparison. L3 bandwidth on EPYC is multi-TB/s.
Fat GPU nodes can readily overload network interfaces so if bisection bandwidth is your concern, CPU nodes are good.
> HPC however, is specifically programmed to be bandwidth-bound instead.
This is wishful thinking. Lots of applications used to justify the US exascale program (and others) are latency-bound. Climate, weather, unsteady CFD, and much of mesoscale materials science and molecular dynamics are run at their latency limit in most scientific studies (one-off scaling studies notwithstanding). There's an unfortunate disconnect between what scientific computing actually needs versus what funders and the media portray.
- dragontamer 6y ago> A64FX is HBM2 at equivalent bandwidth to GPUs (with lower power). CCX is much finer granularity than an entire GPU, so not a direct comparison. But only when using SVE512 SIMD-units, which are grossly similar to GPU SIMD units. At a minimum, SIMD reigns supreme. Even Intel only gets its max bandwidth when using AVX512 units. Once you start rewriting your inner loops to run with SIMD, its not too difficult to start thinking about a dedicated SIMD-accelerator, or GPU, to do the job. > L3 bandwidth on EPYC is multi-TB/s. 32-bytes / infinity fabric cycle. 16-CCX per EPYC chip. 3GHz == 1.5 TBps. I dunno about "multi-TB/s", but its over 1 TBps... sure. But only if used in parallel. EPYC only has 32MBs of L3 per CCX. To achieve the full bandwidth, you need to split the problem into each CCX (which asserts a MESI-like "exclusive" lock when writing to a L3 location: preventing other L3 caches from reading-or-writing there). Even then, EPYC's L3 cache is small compared to GPU-VRAM. Radeon VII 1TBps VRAM applies at full speed with atomics / synchronizations (in fact, the atomics / synchronization to L2 cache. I just don't have L2 numbers for that GPU...). I would expect Radeon VII's L2 cache to have more bandwidth than EPYC's L3 cache (Indeed: the Radeon VII L2 cache is in front of a 1TBps HBM2 cluster). If we traverse up the GPU cache structure, you get 10TBps+ on __shared__ or LDS RAM, which is used as atomic-synchronization points or thread-barriers within a workgroup (a batch of up to 1024 cudaThreads). As such: synchronization between threads (within a large workgroup), or even across the device, is reasonably efficient. The downside of this comparison is that 1024 cudaThreads only have access to 64kB of __shared__ RAM. So it isn't really comparable from a size perspective. ------ Your "TBps" estimate on EPYC's L3 cache however, is misleading. Because you spend significant amounts of MESI messages passing cache lines back and forth between CCX to get there. Compared to a unified L2 on the GPU (or unified VRAM), its just not really comparable. EPYC L3 is somewhere between GPU L2 and GPU __shared__, in terms of memory hierarchy and complexity of use.
- jedbrown 6y agoOne must use SIMD and significant on-node parallelism to saturate a CPU or GPU machine. We have no disagreement there. The question is whether GPUs have a durable architectural advantage for versatile SIMD workloads. Their inverted cache hierarchy and lack of persistence between kernel launches is a concern for memory locality. Regarding L3 bandwidth on EPYC, I have PDE solvers that exceed V100 global memory bandwidth for problem sizes that fit in EPYC L3 cache. I use NPS4 and eschew sharing data structures between CCX. This is the sweet spot for a sizable fraction of apps. The V100 is better value if you have more patience (and are thus able to batch or run larger problem sizes per device).
- dragontamer 6y ago> Their inverted cache hierarchy and lack of persistence between kernel launches is a concern for memory locality. Hmm, I can certainly agree to that. There's some tricks to try to minimize that issue, but they're complicated to use and really screw up the overall architecture. Batching up a larger batch of things to process can help for that. But GPUs have much smaller caches in general and are clearly designed for running out of VRAM. (With caches mainly for atomic / synchronization here and there). A100 only has 40MBs of L2 cache (last level), but the extra threads really eat up that space faster than you expect.