3 ms·
Is the ARM ISA fundamentally faster than x86? If the underlying architectures are converging as described, it kinda seems that the difference between the two in
by pythonvisa 6y ago
Is the ARM ISA fundamentally faster than x86? If the underlying architectures are converging as described, it kinda seems that the difference between the two instruction sets are more legal than technical.
- tambourine_man 6y agoNo, just a lot of brute force and focused optimization going on the M1. There are a few advantages, (it’s a newer architecture after all) but nothing that justifies the current difference.
- toast0 6y agoI think there's a couple things here. x86 defines a pretty strict ordering of memory operations, where ARM is more relaxed; as discussed elsewhere in the thread, this means atomic operations, used for synchronization, will need to wait for pending stores to finish on x86, but not on ARM. The other thing is that the M1 processor seems to be a lot wider than contemporary processors. This means it can (potentially) do more operations per clock. Wider processors are harder to clock faster, but it works for Apple. Apple has no desire to put in the cooling you need to get chips running at 4+ GHz, so it's not a big deal if their chips are clock limited. On the other hand, Intel and AMD like their cores to approach 5 GHz at the top end.
- calo_star 6y agoCan you elaborate on the "Wider processors are harder to clock faster" part?
- normaljoe 6y agoClock is a up volt and down volt. There is a period of time it takes to reach either up or down. If I have to get the clock to reach 2 cores/units at the same time that is easier then say 4 cores/units. As you increase the clock speed the time to reach all the cores decreases. So if you are wider you would need to reach more cores/units in the same period of time, hence "harder". EDIT (To make it a little more clear) To make it more clear the voltage change is typically represented in books as a vertical line, but that is not the case it's diagonal and fuzzy. By fuzzy I mean not a straight line but will have some tiny mini downs on the way up. Different parts of the circuits are going to respond to the up or the down. They are also can vary based on the exact voltage. For example if I have a 5V up one circuit might consider 4.8V to be up and another could be 4.9 or 4.7. Silicon has improved, but there still is limits of scale based on size, volts and timing.
- hehetrthrthrjn 6y agox86, as a legacy architecture, has a lot of baggage and complexity and takes more space and energy to decode and translate instructions to micro-ops. Also because it was not designed as a RISC arch, but rather "backported" as one, it misses out on some RISC advantages. This is why you don't see Intel being competitive in low power applications where ARM excels.
- mhh__ 6y agoLittle yes, little no. The set of x86 instructions people actually use is sufficiently small that the (ignoring the amortized cost of the rest of the ISA - it should be considered, but we can't really know how much space it actually takes up on the die) the two architectures have, to first order at least, converged such that the tricks are now all in the details like memory ordering and scheduling than the actual instructions as per se.
- socialdemocrat 6y agoAs far as I understand, AMD has come out and set it does not make sense for them to make wider than 4 instruction decoders. It seems the CISC architecture creates an upper limit for decoders, as complexity rapidly grows when you have no idea where the next instruction begins in a variable length ISA. So Apple has twice as many decoders, eight, and may actually be able to keep adding to that number while AMD and Intel may get stuck on 4, thanks to the x86 CISC legacy.
- BeeOnRope 6y agoIntel has been at 5 decoders since Skylake (~2015). I'm not sure why everyone focuses on the decoders as the primary determinant of width. There are other bottlenecks which may be narrower than the decoders, and the decoders may not even be used when a uop cache or something like a "loop buffer" (LSD on Intel) is present. So AMD doesn't believe that wider chips aren't useful or that 4 is a limit, because they went to 5-wide in Zen (or 6, depending on how you count it) and I expect them to go wider in the future. Intel went from 4 wide (narrowest bottleneck) to 5 wide in Ice Lake. Wider chips are the future, and there is no "hard wall" at 4 for x86, just like any earlier width increase: just constantly diminishing returns.
- sirn 6y agoIsn't the extra 2 on Zen not a decoder unit, but rather a 6-way dispatch[1] with op-cache being used to kept the execution unit busy? Or was that the reason behind calling it a 5 or 6-wide depending on how one's may count it considering cache miss and all? [1]: https://images.anandtech.com/doci/16214/Zen3_arch_5.jpg https://images.anandtech.com/doci/16214/Zen3_arch_5.jpg