4 ms·
It’s not too dissimilar from how AVX instructions are so poorly implemented on AMD CPUs - they may support it but it may not be core to their strategy (no pun i
by devonkim 8y ago
It’s not too dissimilar from how AVX instructions are so poorly implemented on AMD CPUs - they may support it but it may not be core to their strategy (no pun intended).
- dragontamer 8y agoAVX instructions are actually done fine in my experience. 256-bit is emulated, but that's not really a big deal. Intel is faster for sure, but I never had issues running AVX code on AMD CPUs. The performance characteristics are different between Intel and AMD. In fact, AMD has some tricks up its sleeve. AMD Zen has two AES pipelines (while Intel Skylake only has one), so vectorized AES code is faster on AMD. There are some performance issues with vpgatherdd instructions, but Intel has those as uop emulated code too. So both Intel and AMD are equally to blame there. ---------- My "issues" with AMD CPUs are relatively tame. AMD's profiling tools are weaker than Intel's. (Ex: AMD has "Instruction Based Sampling" while Intel's "PEBS" (Precision Event Based Sampling) are a bit easier to use. True 256-bit execution would be nice, but its not a major hassle in most cases IMO: AMD's back-end is very thick, so you can still get a lot of ILP out of AVX2 instructions.
- floatboth 8y ago> High Precision Event Timers Isn't HPET is just a basic system timer? (That is definitely present on Zen.) You might be thinking of something else? > 256-bit execution would be nice, but its not a major hassle in most cases Coming with Zen 2 anyway :)
- dragontamer 8y ago> Isn't HPET is just a basic system timer? (That is definitely present on Zen.) You might be thinking of something else? You're right. I got the names confused. I've edited the post above to use the proper "PEBS" term. Intel's "Precise Event-Based Sampling" is what I was trying to talk about. Intel's PEBS can precisely tell you where a branch-mispredict happens. AMD's default event timers are inaccurate: your branch mispredictions will be all over "add" instructions, and other unrelated stuff. This is because a CPU is looking at roughly ~100 instruction windows (between the pipeline, decoder, and retirement... there's a lot of inaccuracy in determining "where did this branch misprediction happen??"). So when trying to track down a branch-misprediction on AMD systems, you have to switch to the harder to use "Instruction based Profiling" mode. Intel has a simpler PEBS switch which is easier to use IMO.
- CoolGuySteve 8y agoAt some point I'm hoping to see AVX completely emulated in microcode and replaced with an embedded GPU core that can write to an L3 or L4 cache. This slow scaling of 128->256->512 bits in the instruction set is more or less a solved problem in the GPU space with shader compilers and AVX would mostly be redundant with GPUs if it weren't for the memory bandwidth constraint. ie: When it comes to vector processing, go big or go home. AVX/SSE are a compromise from back when CPU die space and bandwidth was more precious. Now that we have 8-32 cores on a die with a good bus between them, it seems like duplicating those AVX units 8-32 times is less optimal.
- dragontamer 8y agoPlease no. AVX's main advantage is that it is roughly 1-cycle away from your main registers, and 4-cycles away from L1 cache. Talking to and from L3 cache is on the order of 30 to 40 cycles, an order of magnitude slower. If your workload fits inside of 64kB, AVX is incredibly beneficial. If your workload fits within 8MB (L3 cache), you're starting to look at a point where maybe you should pipe that data to the GPU instead. A GPU call over PCIe is under 5 uS / 5000 nanoseconds, with a bandwidth of ~15GB/s. GPUs are certainly farther away than L3 cache, but if you're pushing L3 cache levels... you're getting close to the GPU anyway. --------- 128-bit SSE code is perfect for representing Complex Numbers (Two double-floats). 128-bit is also great for a x,y,z,w 32-bit vector. ------- GPUs are the "go big or go home" architecture. AVX's primary benefit is latency.
- m0zg 8y agoWhat you think of as "vector" processing is currently being used by compilers to speed up things you didn't think were vectorizable. This is possible only because these instructions are pretty cheap latency-wise. By introducing huge latency, you'd be ruining performance of autovectorization, which accounts for a lot of the performance gains in the past decade.
- deleted 8y ago[deleted]
- CoolGuySteve 8y ago