5 ms·
The ARM cortex x4 will have 10-wide dispatch: https://www.anandtech.com/show/18871/arm-unveils-armv92-mobile-architecture-cortex-x4-a720-and-a520-64bit-exclusiv
by celrod 3y ago
The ARM cortex x4 will have 10-wide dispatch:
https://www.anandtech.com/show/18871/arm-unveils-armv92-mobile-architecture-cortex-x4-a720-and-a520-64bit-exclusive/2 https://www.anandtech.com/show/18871/arm-unveils-armv92-mobi...
It'll probably be roughly the same chip as one of their upcoming V chips.
The specs there sound seriously impressive to me, and like they might be getting ready to leave AMD and Intel behind in terms of IPC (for their highest performing chips).
ARM's clock speeds are much lower, so single core performance will probably be worse. But I'd guess server clock speeds may be similar.
- _a_a_a_ 3y agoHaving an N-wide dispatcher means nothing unless the software can use it, and server clock speeds tend to be lower than desktops. Disclaimer: I don't know what I'm talking about
- celrod 3y agoGiven that the M1 and M2 perform similarly to AMD and Intel CPUs with far higher clock speeds, it seems most software can use wider dispatch. Note that dispatch doesn't mean vector width, which is harder for software to take advantage of. It means how many uops the pipeline can handle/clock cycle.
- _a_a_a_ 3y agoA basic block is usually taken to be about 6 instructions. I suppose if you take one or two speculated branches as well then you might easily get to your 10 dispatches possible. Perhaps. As for higher clock speeds, there's a whole lot more that matters such as pipelining instructions, cache sizes, and any number of other things. Clock speed by itself isn't particularly revealing.
- celrod 3y agoFWIW, IIRC my Skylake-X CPU normally has around 2 instructions per clock when I run perf. It has a pipeline width of 4 uops internally. So it's falling far short of a typical basic block/clock cycle. I would also not expect the X4 to utilize that full width in any real workload, but it only needs a fraction of its full width to get more IPC than the X64 competition. But I'd expect it is a reasonably balanced chip (why waste silicon?), and thus the wide pipeline is an indicator of the chip itself being wide with immense out of order capability. Branch prediction rates tend to be extremely high. The X4's frontend also has 10 frontend pipeline stages. Which means ideally, it'd be correctly predicting all branches at least 10 cycles into the future, so that on clock cycle `N-10`, the frontend can get started on the correct instructions that'll be needed on clock cycle `N`. The difference between 1 basic block/cycle and >1 basic block/cycle is really small; it already needs a long history of successful predictions to get 1. But of course, each mispredict is extremely costly. As for bringing up clock speeds and the M1, my point there was that the M1 has already left Intel and AMD behind in terms of IPC; it achieves similar performance despite much lower clock speeds. My original comment said that the ARM Cortex X4 looks like it is starting to leave Intel and AMD behind in terms of IPC, and I used the width as an indicator. You responded saying that the software has to actually allow for this. Yet the M1 example shows that existing software does in fact allow for significantly more out of order execution than Intel and AMD CPUs achieve. So you could argue that, unlike the M1, the Cortex X4 will not be able to realize such an advantage. While plausible, if it does fail to do so, we at least won't be able to blame the software, because the M1 is able to do so despite the software. It'd have to be some deficiency of the X4 relative to the M1 -- such as cache sizes, memory bandwidth... Hopefully it does turn out to be a great chip! But that remains to be seen.
- _a_a_a_ 3y agoI can't imagine getting every instruction in a basic block started at every clock. There is almost certainly dependencies within the block. I talked about basic blocks because that would mean that if you want to kick-off more instructions than in the block, you'd have to speculate about the branch taken. And I am sure can start executing instructions down are speculated jump, only I don't know how far. I also don't know if you can speculate past a write to RAM. I'd like to know. > Yet the M1 example shows that existing software does in fact allow for significantly more out of order execution than Intel and AMD CPUs achieve My point is, shortening the pipeline needed to execute an instruction would also get higher performance. Perhaps they invented a better cache, perhaps larger, perhaps more associative, perhaps…? There's more than one way of increasing performance besides IPC and clock.