5 ms·
I find the vpmullq part the most stunning. This instruction is used in some bignum code, for example if you are implementing RSA. Yet AMD implemented it three
by fefe23 4y ago
I find the vpmullq part the most stunning.
This instruction is used in some bignum code, for example if you are implementing RSA. Yet AMD implemented it three times faster than Intel.
I'm also fascinated by AMD now making AVX512 worthwhile on consumer devices (where they would until quite recently artificially slow down Intel CPUs that had it), which presumably will lead to widespread adoption where it matters. Intels strategy of turning off AVX512 in the recent consumer devices because their energy efficiency cores don't have it may turn out to be a monumental mistake.
- ComputerGuru 4y agoNo one is going to be able to seriously use and support AVX512 (or be sufficiently motivated to implement support for it in their libraries and especially applications) until Intel finally gets its act together with regards to AVX512 and decides it actually wants to commit to it being a thing. The AVX2 rollout was (comparatively) flawless. The gains AVX512 brings over AVX2 are, for most people w/ specialty libs excluded, not worth dealing with the terrible CPU support. And Intel just keeps making the situation worse, taking one step forward and two back.
- jackmott42 4y agoImagine next gen consoles, suppose they stick with AMD. Then every game studio and game engine studio is going to love flinging some AVX-512 around. Developers will get more experience with it, any game that runs on PC and Console is going to look slow on PC if you have intel cpus with bad support. More libraries and tools will get created that people will want to use. Adoption could accelerate quick!
- kllrnohj 4y agoNext-next gen consoles are probably still a good 5+ years away. AVX-512 for consumers will either have already become "a thing" or it'll be dead & buried by then.
- jackmott42 4y agoPeople said that about it 5 years ago to. Yet here we are. Nobody is going to just get rid of it, servers are already using it.
- bayindirh 4y agoThe biggest problem is not support for the instruction set in the silicon, but the performance penalty it brings. SIMD hardware is the most power hungry block on Intel CPUs, and the frequency penalty it brings is never completely disclosed in the tech docs. Even Intel doesn't share that information with you (as a serious customer) sometimes. In HPC world, no instruction is too obscure or niche to use. However, when you use these instructions too frequently, the heat load it generates can slow you down instead of accelerating you over the course of your job, so AVX512 is a pretty mixed case in Intel CPUs. Regardless of this penalty, numeric code benefits from wider SIMD pipelines in most cases. At worst, you see no speedup, but you're investing for the future. On the other hand, we have seen applications which run faster on previous generation hardware due to over-optimization.
- Tuna-Fish 4y ago> However, when you use these instructions too frequently, the heat load it generates can slow you down It's not the heat load that slows you down. If you are using them enough that you produce enough heat that you have to downclock, it's still a win because the instructions improved your throughput more than what you lost in clocks. The problem with Intel's initial AVX-512 implementation was that they didn't clock down because of heat, they clocked down pre-emptively and substantially whenever the CPU executed even a single AVX-512 instruction, even if there was no added heat load, and stayed on the lower clocks for a long period. This worked fine any proper SIMD loads, but was crushing in any situation where there was just a handful of AVX-512 ops between long stretches, such as using an AVX-512 optimized version of some library function.
- bayindirh 4y ago> [T]hey clocked down pre-emptively and substantially whenever the CPU executed even a single AVX-512 instruction... Because you were hitting the power envelope limits in the CPU in these cases too. You might not see the heat, but the CPU cannot carry the power required to keep that core at non-AVX speeds with these power-hungry blocks operated at full speed. As I said, to add insult to the injury, Intel didn't share the exact details of its AVX implementations and frequency ranges it operates, either. Ah, publicly sharing your findings is/was forbidden too.
- mort96 4y agoI don't see why performance-critical code wouldn't have an AVX512 implementation in addition to a scalar or SSE or AVX2 fallback, if AVX512 gives a big enough speed-up on a large enough number of relevant devices.
- janwas 4y ago> The gains AVX512 brings over AVX2 are, for most people w/ specialty libs excluded The last two things I worked on, image compression and quicksort, see 1.4-1.6x end to end speedups from AVX-512 vs AVX2. Is that sufficiently motivating? Especially because the only thing we had to do was ensure that CI machines are AVX-512 capable so that those test codepaths also run. The "terrible CPU support" is a fact of life, not just in x86 (AES is 'optional' in SVE2, sigh), and so we deal with it via runtime dispatch - using what the CPU supports.
- oxxoxoxooo 4y ago> This instruction is used in some bignum code Could you be more specific? I think for that to work one would also need the upper half of 64x64 multiplication and `vpmullq` provides only the lower half. You could break one 64x64 multiplication into four 32x32 multiplications (i.e. emulate the full 64x64 = 128 bits multiplication) but I was under the impression that this was slow.
- adrian_b 4y agoI assume that as you say, whoever used this instruction was using it for multiplying 32-bit numbers. On AMD Zen 4 and Intel Cannon Lake or newer (when AVX-512 is supported), the fastest method to multiply big numbers is to use the IFMA instructions, which reuse the floating-point multipliers to generate 104-bit products of 52-bit numbers.
- pbsd 4y agovpmullq is not that useful; in bignum code you also want the upper part of the product, and there is no corresponding vpmulhq instruction to get that. On the other hand, vpmadd52luq and vpmadd52huq do give you access to the lower and upper parts of a 52x52->104 bit product, and those instructions perform well in the Intel chips, 3x faster than vpmullq.