3 ms·
The "wall of text" TL;DR: - +16 registers (thus 32) and optionally separate destination (looking very RISC like now) - PUSH2/POP2 with full forwarding - Much
by FullyFunctional 3y ago
The "wall of text" TL;DR:
- +16 registers (thus 32) and optionally separate destination (looking very RISC like now)
- PUSH2/POP2 with full forwarding
- Much expanded predication, including predicated loads and stores
This is pretty interesting. Especially the latter can make a big difference for highly unpredictable memory intensive code, like compression.
- _old_dude_ 3y agoAlso AVX-10, E-cores and P-cores having 512 bits vector registers (see the last references).
- jcranmer 3y agoIf I read the note correctly, P-cores won't have 512-bit vector registers, but they will have the other fancy stuff added by AVX-512 (namely, vector predication stuff, static rounding mode instructions, new vector instructions like complex multiply or half-precision float, and 32 vector registers), just only for 128-bit and 256-bit vectors. Which, to be fair, is arguably the more useful parts of AVX-512 anyways; the maximum vector length being upped isn't all that interesting.
- coder543 3y ago> If I read the note correctly, P-cores won't have 512-bit vector registers I don’t entirely agree. See the second graphic here: https://www.phoronix.com/news/Intel-AVX10 https://www.phoronix.com/news/Intel-AVX10 Mentioned above the graphic: > Part of making AVX10 suitable for both P and E cores is that the converged version has a maximum vector length of 256-bits and found with the E cores while P cores will have optional 512-bit vector use. 512-bit support will be optional, so maybe every P-core won’t have it… maybe it’ll be restricted to higher end processors? But it sounds like some will have it, or it wouldn’t be an option at all.
- adgjlsfhk1 3y agopresumably it will be the server chips that are all p cores that have it. to me it seems like a dumb choice, but it is consistent with what Intel is doing now
- wmf 3y agoP-cores will be 512-bit and E-cores will be 256-bit. Seems unnecessarily complex to me.
- adrian_b 3y agoNo, see another more complete reply above. P-cores in hybrid CPUs will be 256-bit (like E-cores). P-cores in (server) CPUs that contain only P-cores will be 512-bit.
- _old_dude_ 3y agoyes, thanks,
- deleted 3y ago[deleted]
- adrian_b 3y agoWhat Intel says exactly: "A “converged” version of Intel AVX10 with maximum vector lengths of 256 bits and 32-bit opmask registers will be supported across all Intel processors, while 512-bit vector registers and 64-bit opmasks will continue to be supported on some P-core processors." So all future Intel CPUs starting in 2025 will support a 256-bit subset of AVX-512, where AVX-512 is rebranded as AVX10. Only some P-core processors will support the full 512-bit AVX-512 a.k.a. AVX10, which is to be understood that only those server CPUs that contain only P-cores, i.e. the successors of Granite Rapids and Granite Rapids D, will support 512-bit registers and instructions (and 64-bit mask registers instead of 32-bit mask registers).
- adrian_b 3y agoNo, the E-cores will implement only a 256-bit subset of AVX-512, which halves the size of the vector registers to 256-bit and the size of the mask registers to 32-bit. The same subset will be implemented on the P-cores combined with E-cores. This subset AVX10/256, is the reason for this new specification. It is the Intel response to AMD Zen 4. When their competitor supports AVX-512 on all products, Intel had to do something to remain competitive. Because they believe that supporting the full AVX-512 on their E-cores is too expensive, they have created a subset of AVX-512, including only the instructions with an operand size up to 256 bits.
- Bulat_Ziganshin 3y ago128/256-bit subset of AVX-512 already exists and implemented by P-cores as well as Zen4. They will just enable it via micro-code.
- adrian_b 3y agoEven if since Skylake server the AVX-512 ISA includes scalar, 128-bit vector, 256-bit vector and 512-bit vector instructions, it was not possible to implement any subset that did not include the 512-bit vector instructions, because there were no means for a program to discover that the 512-bit instructions are missing. Now, a different method has been defined for discovering through CPUID which AVX-512 a.k.a. AVX10 features are implemented, so only now it has become possible to implement an up to 256-bit subset. Moreover when 128-bit vector and 256-bit vector instructions have been added to AVX-512, 2 bits from the EVEX prefix that were previously used for rounding control have been reused to encode the length of the vector operands. Because of this, only the 512-bit vector instructions and the scalar instructions can specify the rounding control. So if the 512-bit vector instructions are deleted, there is no longer any way to specify the rounding control for vector instructions. To solve this problem, in the first CPUs that will implement the 256-bit subset of AVX-512, new encodings will be used for the 256-bit instructions with rounding control. Also the XSAVE and XRSTOR instructions had to be modified to save and restore correctly the new vector registers and mask registers. So implementing the 256-bit subset of a AVX-512, a.k.a. AVX10/256, is not so straightforward as a microcode update, it requires changes in the instruction decoders and in the structure of the CPUID registers and other smaller changes.
- peterfirefly 3y ago+ Two-byte REX2 prefix that replaces the REX prefix (and 0F prefix). REX2 is D5 + a byte that is a lot like the lower nibble of a REX prefix, only twice as big. It has two extra bits for the up to three registers that can be named in normal x86 instructions: either two registers or a register operand and a memory operand that can use two registers + a displacement for the memory address. It also contains the W bit like REX does (64-bit). It also has a bit called M0 that indicates whether the instruction is in opcode map 0 (primary opcode map) or opcode map 1 (0F opcode map). That means that a REX2 instruction from opcode map 1 takes the same number of bytes as a REX instruction from opcode map 1. Instructions from opcode map 0 are one byte longer with REX2 than with REX. Some of the normal ALU instructions also get an EVEX encoding (in map 4). That allows for a different data destination than before (separate from source the operand(s)). It also allows for ALU instructions that don't change the flags, which must be really nice for the out of order/data forwarding circuitry.
- drudru 3y ago0xD5 was the old BCD instruction: AAD