4 ms·
Good analysis. It's also worth pointing out that this is for 2x 512 bit FMA, which is more than client Ice/Tiger/Rocket lake or Zen4 have. Personally, I bought
by celrod 4y ago
Good analysis. It's also worth pointing out that this is for 2x 512 bit FMA, which is more than client Ice/Tiger/Rocket lake or Zen4 have.
Personally, I bought HEDT (Skylake-X and Cascadelake) because I wanted 2x 512 bit AVX512. Glad it's cheap in terms of area, and I'm hoping we'll get more options with great vector performance in the future.
- janwas 4y agoGood point about the second FMA. I'm not certain it's the best tradeoff, Genoa only has two half-width FMA. I share your hope for more focus on vectors. It's also up to us software devs, CPUs will not invest as heavily if we don't use it.
- celrod 4y agoI do think that Genoa's approach is a reasonable one. I'd like to see one of Gracemont's successors doing the same. Maybe we'll even see quadruple pumping for AVX-512 some day? I'll be impressed if/when an Atom line CPU gets 4x 128bit fma units to match ARM's Cortex-X line or Apple's Firestorm). I think these are good options, and can allow AVX-512 to sort of act like SVE, but with the benefits of a fixed size architecture (i.e., shuffles); compile one set of code and you're able to run it anywhere, with performance dictated by how much the vendor decided was worth investing into the vector units. And AVX512 can still help (like it does Genoa) by taking a lot of pressure off of the front end. I'm also trying to do my part as a software dev! I wrote/maintain the JuliaSIMD ecosystem, and working on good loop vectorizers to let people take advantage of their vector units is my passion; LoopVectorization.jl has gotten great results on many benchmarks[0], and I'm rewriting it as an LLVM pass to try and address as many of its flaws and limitations as I can. [0] For example, in a simple self-dot-product benchmark, LLVM's 256 bit code is actually faster than its 512 bit code when testing random sizes from 1-256: https://github.com/JuliaSIMD/LoopVectorization.jl/issues/446#issuecomment-1331500465 https://github.com/JuliaSIMD/LoopVectorization.jl/issues/446... However, LoopVectorization.jl's 512 bit code is close to twice as fast as either LLVM's 256 bit or 512 bit code. This is a trivial example; the difference can be much larger for more complicated code.
- janwas 4y ago> Maybe we'll even see quadruple pumping for AVX-512 some day? > can allow AVX-512 to sort of act like SVE, but with the benefits of a fixed size architecture (i.e., shuffles); compile one set of code and you're able to run it anywhere, with performance dictated by how much the vendor decided was worth investing into the vector units That makes a lot of sense. It's basically the equivalent of RISC-V's LMUL=4, with the big advantage of reducing instruction count as you say. That seems a better route than 4x128, which might actually be less in practice if there are resource conflicts. > I'm also trying to do my part as a software dev! I wrote/maintain the JuliaSIMD ecosystem That's awesome, congrats on the good result. Looks like your preference is to allow people to write high-level code without much worry about the arch details. Any thoughts on how we can spread awareness of the basics such as data-oriented programming (avoiding branches, optimizing for cache and contiguous memory accesses)?