3 ms·
One must use SIMD and significant on-node parallelism to saturate a CPU or GPU machine. We have no disagreement there. The question is whether GPUs have a durab
by jedbrown 6y ago
One must use SIMD and significant on-node parallelism to saturate a CPU or GPU machine. We have no disagreement there. The question is whether GPUs have a durable architectural advantage for versatile SIMD workloads. Their inverted cache hierarchy and lack of persistence between kernel launches is a concern for memory locality.
Regarding L3 bandwidth on EPYC, I have PDE solvers that exceed V100 global memory bandwidth for problem sizes that fit in EPYC L3 cache. I use NPS4 and eschew sharing data structures between CCX. This is the sweet spot for a sizable fraction of apps. The V100 is better value if you have more patience (and are thus able to batch or run larger problem sizes per device).
- dragontamer 6y ago> Their inverted cache hierarchy and lack of persistence between kernel launches is a concern for memory locality. Hmm, I can certainly agree to that. There's some tricks to try to minimize that issue, but they're complicated to use and really screw up the overall architecture. Batching up a larger batch of things to process can help for that. But GPUs have much smaller caches in general and are clearly designed for running out of VRAM. (With caches mainly for atomic / synchronization here and there). A100 only has 40MBs of L2 cache (last level), but the extra threads really eat up that space faster than you expect.