32 ms·
Nice to see that AVX512 hasn't died with Xeon Phi. I see it coming out in a number of high end but lightweight notebooks too (Surface Pro with i7 10XXG7, MacBo
by marmaduke 6y ago
Nice to see that AVX512 hasn't died with Xeon Phi. I see it coming out in a number of high end but lightweight notebooks too (Surface Pro with i7 10XXG7, MacBookPro 13" idem). This is a nice way to avoid needing GPU for heavily vectorizable compute tasks, assuming you don't need the CUDA ecosystem.
- 37ef_ced3 6y agoFor example, AVX-512 neural net inference: https://NN-512.com https://NN-512.com Only interesting if you care about price (dollars spent per inference) For raw speed (no matter the price) the GPU wins
- dragontamer 6y agoGPGPU will never really be able to take over CPU-based SIMD. GPUs have far more bandwidth, but CPUs beat them in latency. Being able to AVX512 your L1 cached data for a memcpy will always be superior to passing data to the GPU. With Ice Lake's 1MB L2 cache, pretty much all tasks smaller than 1MB will be superior in AVX512 rather than sending it to a GPU. Sorting 250,000 Float32 elements? Better to SIMD Bitonic sort / SIMD Mergepath (https://web.cs.ucdavis.edu/~amenta/f15/GPUmp.pdf https://web.cs.ucdavis.edu/~amenta/f15/GPUmp.pdf) on your AVX512 rather than spend a 5us PCIe 4.0 traversal to the GPU. It is better to keep the data hot in your L2 / L3 cache, rather than pipe it to a remote computer (even if the 16x PCIe 4.0 pipe is 32GB/s and the HBM2 RAM is high bandwidth once it gets there). -------- But similarly: CPU SIMD can never compete against GPGPUs at what they do. GPUs have access to 8GBs @500GB/s VRAM on the low-end and 40GBs @1000GB/s on the high end (NVidia's A100). EDIT: Some responses have reminded me about the 80GB @ 2000GB/s models NVidia recently released. CPUs barely scratch 200GB/s on the high end, since DDR4 is just slower than GPU-RAM. For any problem where data-bandwidth and parallelism is the bottleneck, that fits inside of GPU-VRAM (such as many-many sequences of large scale matrix multiplications), it will pretty much always be better to compute that sort of thing on a GPU.
- celrod 6y agoFWIW, the A64FX has 1TB/s bandwidth because it has 32GiB of HBM2.
- dragontamer 6y agoAh yes, good point. ------- It should be noted that most CPUs do use DDR4 for good reason: they have an expectation to use more than the capacity of HBM2. GPUs (and a supercomputer CPU, like the A64FX) can rely upon the fact that their workloads are designed for distributed-compute and/or otherwise fit inside of 32GBs or 40GBs. CPUs on the other hand, could very well be reading/writing to a 2TB in-memory database these days, even on a single-socket single-node system like AMD EPYC. That flexibility to have as much (or as little) RAM as the customer demands is a major advantage of traditional CPUs like EPYC or Xeon, which the A64FX cannot partake in. If you need more than 32GBs of RAM, A64Fx is a no-go. That HBM2 is soldered directly onto the package and cannot expand.
- celrod 6y agoYes, and 32 GB seems rather limited for a 48 core chip. While not a "general purpose" CPU, this HPC CPU does show IMO that it's at least possible of making a CPU that does compete in some aspects with what GPUs do. Although an A100 still has twice the bandwidth and over 3x the GFLOPS without tensor cores and more than 6x with.
- ajross 6y agoFWIW: your DRAM numbers are quoting clock speeds and not bandwidth. They aren't linear at all. In fact with enough cores you can easily saturate memory that wide, and CPUs are getting wider just as fast as GPUs are. The giant Epyc AMD pushed out last fall has 8 (!) 64 bit DRAM channels, where IIRC the biggest NVIDIA part is still at 6.
- dragontamer 6y ago> 8 (!) 64 bit DRAM channels Yeah. And at 3200 Mbit/sec, that comes out to 200GB/s. (3200 MHz x 8-bytes (aka 64-bit) == 25GB/s. x8 channels == 200GB/s). > where IIRC the biggest NVIDIA part is still at 6. That's 6x *1024-bit* HBM2 channels. Total bandwidth is 2000GBps, or over 10x the speed of the "8x channel EPYC". Yeah, HBM2 is fat, extremely fat. ---------- *ONE* HBM2 channel offers over 300GBps bandwidth. And the A100 has *SIX* of them. Literally ONE HBM2 channel beats the speed of all 8x DDR4 EPYC memory channels working in parallel.
- ajross 6y agoYou're still quoting clock speeds. That's not how this works. Go check a timing diagram for a DRAM cycle in your part of choice and do the math.
- dragontamer 6y agoDo you know what 3200MHz / PC4-25600 DDR4 means? 25600 is the channel rate in (EDIT) MB/sec of the stick of RAM. That's 25GB/s for a 3200 MHz DDR4 stick. x8 (for 8-channels working in parallel) is 200GB/s. ----------- This has been measured in practice by Netflix: https://2019.eurobsdcon.org/slides/NUMA%20Optimizations%20in%20the%20FreeBSD%20Network%20Stack%20-%20Drew%20Gallatin.pdf https://2019.eurobsdcon.org/slides/NUMA%20Optimizations%20in... As you can see, Netflix's FreeBSD optimizations have allowed EPYC to reach 194GB/s measured performance (or just under the 200GB/s theoretical). And only with VERY careful NUMA-tuning and extreme optimizations were they able to get there.
- 6y ago
- volta83 6y ago> Being able to AVX512 your L1 cached data for a memcpy will always be superior to passing data to the GPU. The two last apps I worked on have been GPU-only. The CPU process starts running and launches GPU work, and that's it, the GPU does all the work until the process exits. There is no need to "pass data to the GPU" because data is never on CPU memory, so there is nothing to pass from there. All network and file I/O goes directly to the GPU. Once all your software runs on the GPU, passing data to the CPU for some small task doesn't make much sense either.
- dragontamer 6y agoSo we know that GPUs are really good at raytracing and matrix multiplication, two things that are needed for graphics programming. However, the famous "Moana" scene for Disney-level productions is a 93GB (!!!!) scene statically, with another 131GBs (!!!) of animation data (trees blowing in the winds, waves moving on the shore, etc. etc.). That's simply never going to fit on a 8GB, 40GB, or even 80GB high-end GPU. The only way to work with that kind of data is to think about how to split it up, and have the CPU store lots of the data, while the GPU processes pieces of the data in parallel. https://www.render-blog.com/2020/10/03/gpu-motunui/ https://www.render-blog.com/2020/10/03/gpu-motunui/ Which has been done before, mind you. But it should be noted that the discussion point for GPU-scale compute runs into practical RAM-capacity constraints today, even on movie-scale problems from 5 years ago (Moana was released in 2016, and had to be rendered on hardware years older than 2016). Moana scene is here if you're curious: https://www.disneyanimation.com/resources/moana-island-scene/ https://www.disneyanimation.com/resources/moana-island-scene... ---------- But yes, if your data fits within the 8GBs GPU (or you can afford a 40GB or 80GB VRAM GPU and your data fits in that), doing everything on the GPU is absolutely an option.
- oivey 6y agoWe know that GPUs are really good at far more than ray tracing and matrix multiplication. Oversimplifying a bit, they’re great at basically any massively parallel operation that has minimal branching and can fit in memory. Using a GPU to just add two images together probably isn’t worth it, but many real world workflows allow you to operate solely on the GPU. If you’re Disney, you can afford boxes with 10+ A100s with NVLink sharing the memory in a single 400+ GB pool. Unknown if that ends up being more economical than the equivalent CPU version, but it’s important to understand in order to evaluate the future of GPUs.
- marmaduke 6y agoIn my experience, the most important aspect missing in most CPU GPU discussions, is that CPUs have a massive cache compared to GPUs, and that cache has pretty good bandwidth (~30 GB/core?), even if main memory doesn't. So even if your task's hot data doesn't fit in L2 but in L3/core, AVX-whatever per core processing is a good bet regardless of what a GPU can do. Another aspect that seems like a hidden assumption in CPU-GPU discussions is that you have the time-energy-expertise budget to (re)build your application to fit GPUs.
- dragontamer 6y agoOn the memory perspective, I basically see problems in roughly the following grouping of categories: 40TBs+ -- Storage-only solutions. "External Tape Merge sort algorithm", "Sequential Table Scan", etc. etc. (SSDs or even Hard drives if you go big enough) 4TB to 40TBs -- Multi-socket DDR4 RAM is king (8-way Ice Lake Xeon Scalable Platinum will probably reach 40TBs). Single-node distributed memory with NUMA / UPI to scale. 1TB to 4TB -- Single Socket DDR4 RAM (EPYC, even if at 4x NUMA. Or Single-node Ice Lake). 80GB to 1TB -- DGX / NVlink distributed memory A100 ganging up HBM2 together. GPU-distributed RAM is king. 256MBs to 80GBs -- HBM2 / GDDR6 Graphics RAM is king (80GB A100 2TB/s). 1.5MBs to 256MBs -- L3 cache is king (8x32MBs EPYC L3 cache, or POWER9 110MB+ L3 cache unified) 128kB to 1.5MBs -- L2 cache is king (1.25MB Ice Lake Xeons L2, this article) 1kB to 128kB -- L1 cache is king. (128kB L1 cache on Apple M1). Note: "GPU __Shared__" is a close analog to L1 and competes against it, but is shared between 32 to 256 GPU threads, so its not an apples-to-apples comparison. 1kB and below -- The realm of register-space solutions. (See 64-bit chess engine bitboards and the like). Almost fully CPU-constrained / GPU-constrained programming. 256x 32-bit GPU registers per GPU-thread / SIMD thread. CPUs have fewer nominal registers, but many "out of order" buffers or "reorder buffers" that practically count as register storage in a practical / pragmatic sense. CPUs just use their "real registers" as a mechanism to automatically discover parallelism in otherwise single-thread written code. ------------ As you can see: GPUs win in some categories, but CPUs win in others. And these numbers change every few months as a new CPU and/or GPU comes out. And at the lowest levels: CPUs and GPUs cannot be compared due to fundamental differences in architecture. For example: GPU __shared__ memory has gather/scatter capabilities (the NVidia PTX instructions / AMD GCN instructions permute vs bpermute), while CPUs traditionally only accelerate gather capabilities (pshufb), and leave vgather/vscatter instructions to the L1 cache instead. GPUs have 32x ports to __shared__, so every one of the 32-threads in a wave-front can read/write every single clock-tick (as long as all 32 they are on different ports/alignment, or you have a special one-to-all broadcast). CPUs only have 2 or 4 ports, so vscatter and vgather operate slowly, as if a single thread were reading/writing each of the memory locations. But CPU L1 cache has store-forwarding, MESI + cache coherence, and other acceleration features that GPUs don't have. GPUs are therefore more efficient at sharing data within workgroups of ~256 threads, but CPUs are more efficient at sharing data between cores, or even among out-of-die NUMA solutions, thanks to robust MESI messaging.
- api 6y agoThe 2020 Intel MacBook Air and 13" Pro have 10nm Ice Lake with AVX512. The Ice Lake MacBook Air performs pretty well and very close to the Ice Lake Pro, though of course the M1 destroys it.
- bitcharmer 6y agoAVX-512 is an abomination in my field and we avoid it like the plague. It looks like we're not the only ones. Linus has a lot to say about it as well. https://www.phoronix.com/scan.php?page=news_item&px=Linus-Torvalds-On-AVX-512 https://www.phoronix.com/scan.php?page=news_item&px=Linus-To...
- aseipp 6y agoSkylake-X has already had its die shots examined and the AVX-512 register file, the dominant part of the layout, is something like .5 of a single core, so it's not even going to buy you much area for anything if Intel deleted it, the whining about how it's better spent on extra cores by Linus is totally overblown. Ice Lake has also dramatically improved the per-core frequency tuning for client SKUs to the point AVX-512 is quite viable on my laptop with no serious problems; a single thread doing something isn't going to tank anything. Ice Lake-X almost certainly has 2FMAs instead of the 1FMA of client SKUs however, so it'll be interesting to see what the new power licensing situation is, but this is clearly something they've had in the books to improve. The problem is that for the workloads that need specialization, you sometimes really need it. You could also delete the vectorized AES units in your Intel machines too and the general purpose performance wouldn't be affected much, but cryptographic performance specifically would tank, and it turns out, that matters a lot in aggregate for many people. Ultimately there are literally dozens of specialized inactive units on any CPU at any given time that could be "better spent on general purpose units" (which also isn't necessarily true if other architectural choices prevent those units from being utilized effectively). People just like complaining about AVX-512 because it's easily digestible water cooler chat they read about on a blog.
- bitcharmer 6y agoAVX-512 is problematic in multiple ways your comment fails to even touch upon. It impedes latency-sensitive applications due to forced down-clocking taking place for AVX-intensive instruction streams. The other side effect is that it quickly produces extreme heat that not only forces further down-clocking of the core but also taints neighbouring cores with dissipated heat and prevents them from going into turbo. Not everything is about performance in laptops.