3 ms·
I'd say the ways to increase performance are 1. decrease the number of instructions needed (has to be done in software, but also dependent on ISA, e.g. using AV
by celrod 3y ago
I'd say the ways to increase performance are
1. decrease the number of instructions needed (has to be done in software, but also dependent on ISA, e.g. using AVX512 can help a lot here, so long as you don't end up executing more scalar epilogue iterations).
2. increase IPC (obviously software can help a lot here)
3. increase clocks (not much software can do here; wider instructions are generally worth it, so if choosing between "1." and "3." in software, it's generally better to favor "1." (especially on more recent CPUs that don't have downclocking problems).
Design of the CPU can also influence all three of these.
Things like better cache, better branch predictors and shorter pipelines, will all help IPC.
> My point is, shortening the pipeline needed to execute an instruction would also get higher performance.
This wouldn't increase throughput if 100% of branches are predicted correctly -- except for the the extra cycles before instructions start executing.
It'd decrease branch mispredict penalties though, which is a big deal and would help in practice. The Cortex X4 did shave off a frontend pipeline stage relative to the Cortex X3 (11 -> 10). This is better than Intel Alder Lake. One contributor is probably that it is easier to decode ARM instructions in parallel without needing mutliple pipeline stages (e.g., one to find out where variable width instructions end before the instruction byte stream can be sent to decoders [not an issue if instructions are already in the uop cache]).
> There's more than one way of increasing performance besides IPC and clock.
I assume by "IPC" here you mean dispatch width?
Things like a better cache for fewer misses, better prefetching, better branch prediction, larger reorder buffers so that it can speculate further ahead before stalling, all help IPC.
Zen1 CPUs (6 uops) were already wider than Intel Skylake (4 uops) and Ice/Tiger lake (5 uops), matching Alder Lake (6 uops).
But they were obviously far behind in IPC (and Zen1 in particular also decoded AVX2 instructions into 2 uops).
Zen1 has SMT, which was part of the reason to go wide early on: the frontend wasn't good enough to feed that width with a single thread, but using two threads could mitigate that. Early on, Zen1 (and the Zen family) generally did better in multithreaded than single threaded benchmarks thanks to that approach.
The ARM Cortex X4 doesn't have SMT, so it's taking a different approach to performance.
A single number isn't going to be representative of performance across benchmarks or all the tasks you're interested in.
Unfortunately, I think it'll be more than a year before we can see the Cortex X4 (as it's aiming at TSMC N3E), but I'm definitely looking forward to deep dives into it's performance (and also that of Intel's Meteor Lake, Zen5, etc).