19 ms·
Unlike what has been said on Twitter the answer to why the M1 is fast isn’t due to technical tricks, but due to Apple throwing a lot of hardware at the problem.
by benjaminl 6y ago
Unlike what has been said on Twitter the answer to why the M1 is fast isn’t due to technical tricks, but due to Apple throwing a lot of hardware at the problem.
The M1 is really wide (8 wide decode) and has a lot of execution units. It has a huge 630 deep reorder buffer to keep them all filled, multiple large caches and a lot of memory bandwidth.
It is just a monster of a chip, well designed balanced and executed.
BTW this isn’t really new. Apple has been making incremental progress year by year on these processor for their A-series chips. Just nobody believed those Geekbench benchmarks showing that in short benchmarks your phone could be faster than your laptop. Well turns out that given the right cooling solution those benchmarks were accurate.
Anandtech has a wonderful deep dive into the processor architecture.
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-deep-dive/2 https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
Edit: I didn’t mean to disparage Apple or the M1 by saying that Apple threw hardware at the problem. That Apple was able to keep power low with such a wide chip is extremely impressive and speaks to how finely tuned the chip is. I was trying to say that Apple got the results they did the hard way by advancing every aspect of the chip.
- vmchale 6y ago> Just nobody believed those Geekbench benchmarks showing that in short benchmarks your phone could be faster than your laptop. I saw a paper on (I think?) SMT solvers on iPhone. Turned out to be faster than laptops, I kind of brushed over it as irrelevant at the time.
- FpUser 6y agoI did believe those benchmarks. I also knew that sustainable load at those speeds if not throttled down would just melt the thing so yes they were irrelevant.
- simonh 6y agoMaybe in practice at the time, but in hindsight they were actually a valid indication of the true performance capabilities of the architecture. It’s just that a phone has too little thermal capacity to sustain the workload.
- greyhair 6y agoAs well as too little battery capacity to sustain the workload. Applying those cell phone methods to a laptop with better cooling and a larger battery was a win.
- FpUser 6y ago>"Maybe in practice at the time" Exactly this. If I am shopping now I do not care how particular CPU/architecture evolves in the future. I only care about what it can practically do now and at what price. As it is now M1 has 4 fast cores and not upgradeable maximum 16GB. For many people it would be more than they ever need. Let them be happy. For my purposes I am way more happy with 16 core new AMD coupled with 128GB RAM (my main desktop at the moment). It runs at sustainable 4.2 GHz without any signs of thermal throttling. Cooler is keeping it at 60C.
- fiddlerwoaroof 6y agohttps://www.cs.utexas.edu/~bornholt/post/z3-iphone.html https://www.cs.utexas.edu/~bornholt/post/z3-iphone.html
- Kolja 6y ago> Just nobody believed those Geekbench benchmarks showing that in short benchmarks your phone could be faster than your laptop. Except for a lot of the Apple-centric journalists and podcasters, who have been imagining for years how fast a desktop built on these already-very-fast-when-passively-cooled chips could be. Not that that matters very much when experienced and real-world workload performance suffers, but as far as I can tell, the M1 is no slouch in that respect either.
- fiddlerwoaroof 6y agoYeah, Apple’s CPUs have been doing really well for a while now: https://www.cs.utexas.edu/~bornholt/post/z3-iphone.html https://www.cs.utexas.edu/~bornholt/post/z3-iphone.html
- pletnes 6y agoI heard 10 years ago, or whatever, that the ipad 2 had the most power efficient CPU available, period. This was told at a keynote by an HPC scientist who cannot be said to be a «apple journalist». Apple have been doing well for a very long time and I’ve expected this moment since that keynote, basically.
- Kolja 6y agoOK, I misspoke. I heard the sentiment from the Apple-focused voices in my information bubble, doesn't mean that nobody else said it. It's just that "nobody believed those Geekbench benchmarks" is not completely true.
- saberience 6y agoI find it interesting you use this kind of disparaging tone when discussing Apple Silicon. I also find it interesting that you consider having a wide decoder not as a technical trick but as "throwing hardware at the problem." However you try and spin it, what it comes down to is this, Apple is somehow designing more performant processors than every other company in the world and we should acknowledge they are beating the "traditional" chip designing companies handily while being new at the game. If it's as easy as "throwing hardware" at the problem, then Intel and AMD and Samsung etc should have no problem beating Apple right?
- robert_foss 6y agoApple is throwing money at the problem. IC size is directly proportional to cost. They couldn't sell a chip like this at a cost-competitive price on the open market against AMD/Intel products.
- libria 6y agoJust to quantify your adjectives, per the Anandtech article: > The M1 is really wide (8 wide decode) In contrast to x86 CPUs which are 4 wide decode. > It has a huge 630 deep reorder buffer By comparison, Intel Sunny/Willow has 352.
- deleted 6y ago[deleted]
- natchy 6y agoSo Intel and AMD are capable of building a chip like this, but the ambitious size meant it was more economically feasible for Apple to build it themselves? (Not a hardware guy)
- ArchOversight 6y agoRead the article linked, it explains why Intel and AMD are unable to throw more decoders at the problem.
- outworlder 6y agoMaybe they are, assuming there's sufficient area in the die for this. They would likely still be massive power hogs.
- masklinn 6y agoIt's not that it was more economical, but that at least some of these AMD and Intel would not benefit from due to the ISA: x64 instructions can be up to 15 bytes, so just finding 8 instructions to decode would be costly, and I assume Intel and AMD think more so than the gains from more decoders (you couldn't keep them fed enough to be worth it, basically).
- SlipperySlope 6y agoWhen ARM migrated to 64 bits, wisely they left the 32 bit instructions behind. So easy to decode and uniform instruction length. On the other hand, AMD64 can run 8088 16 bit instructions, which means that it must handle a variety of instruction lengths.
- sliken 6y agoYou mention most of the big changes, except one. Assuming a random (but TLB friendly) pattern the M1 manages a latency of around 30-33ns to main memory, about half of what I've seen anywhere else. Impressive. Maybe motherboards should stop coming with dimms and use the apple approach to get great bandwidth and latency and come in 16, 32, and 64GB varieties by soldering LPDDR4x on the motherboard.
- brandmeyer 6y ago> Assuming a random (but TLB friendly) pattern the M1 manages a latency of around 30-33ns to main memory. This, right here. It also helps that the L1D is a whopping 128 kB and only 3 cycles of load-use latency.
- AlphaSite 6y agoI wonder how they managed that.
- my123 6y agoThe NVIDIA Carmel cores on 12nm had a 64KB L1D cache with a 2 cycles latency.
- throwaway_pdp09 6y agoMeans nothing without saying what the clock goes at.
- my123 6y ago2.26GHz, on a quite old process.
- xpuente 6y agoHuge block size (128bytes). Probably they are using Power7 alike scheduling (i.e. scheduling are working on packs of instructions, That might explain the humorous 600+ entry ROB. Certainly the wake-up logic can't deal with that one-by-one with such a low power). If you combine that with JIT and/or good compilers, you get this. I guess only Apple can pull this trick: they control all the stack (and some key power architects are working there).
- ori_b 6y agoThe next question: What prevents Intel or AMD from doing this on their processors?
- jonplackett 6y agoThe article specifically answers this: - x86 instruction set can't be queued up as easily because instructions have different lengths - 4 decoders max, while Apple has 8 and could go higher. - Business model does not allow this kind of integration.
- masklinn 6y ago> instructions have different lengths also allows extremely long instructions, the ISA will allow up to 15 bytes, and fault at 16 (without that artificial limit you can create arbitrarily long x86 instructions).
- sterlind 6y agoWhat a nightmare, but it makes me wonder: rather than decoding into micro-ops at runtime, could Intel or AMD "JIT" code up-front, in hardware, into a better bytecode? I'm sure it wouldn't work for everything, but why wouldn't it be feasible to keep a cache of decoding results by page or something?
- oseityphelysiol 6y agoFrom what I understand this is exactly what the instruction decoder does.
- annilt 6y agoThey do something similar for ‘loops’. CPU doesn’t decode same instructions over and over again, just using them from ‘decoded instruction cache’ which has capacity around 1500 bytes.
- djcapelis 6y agoThis is exactly how the hardware works and what micro-ops are, on any system with a u-op cache or trace cache those decoded instructions are cached and used instead of decoding again. Unfortunately you still have to decode the instructions at least once first and that bottleneck is the one being discussed here. This is all transparent to the OS and not visible outside a low level instruction cache though, which means you don’t need a major OS change, but arguably if you were willing to take that hit you could go further here.
- gigatexal 6y agoYeah the article is interesting I just like knowing that Apple will keep iterating and the performance gap between Apple Silicon and x86 will continue to grow. I keep spec'ing out an Apple M1 Mac mini only to not pull the trigger because I am curious what an M2 will hold.
- libria 6y agoHah, won't this perpetuate? Whenever M(N) is released M(N+1) will be on the horizon with even greater promise.
- Moru 6y agoYes, been doing that for about 30 years. Now and then you just have to take a leap. I usually buy what was the best last year at a bargain price instead.
- megameter 6y agoThis is a good year for getting last year's stuff given how the supply is so low on most of the high profile hardware(new Ryzens, current generation dedicated graphics, game consoles). I did get something launched this year that reviewed well - a gaming laptop(Legion 5) - but it was not hard to find and I even got it used like-new. Perhaps because nobody's travelling now.
- 2muchcoffeeman 6y agoI think they old adage, “don’t buy first generation Apple products” applies here. Seems sensible to wait to see where this goes.
- gigatexal 6y agoYes true but as the respondent said above I’m just waiting for the software to catch up and maybe a redesign on the laptop side. In the mean time I’m saving for one.
- 6y ago
- PeterisP 6y agoAre they really throwing more hardware at the problem? The die size for the whole M1 SoC is comparable to or even smaller than Intel processors, and the vast majority of that SoC is non-CPU related stuff, the CPU cores/cache/etc seem to be at most 20% of that die - though a more dense die because of the 5nm process. This also seems to imply that the 'budget' of number of transistors for that CPU-part of the SoC is also comparable to previous Intel processors, not a significant increase. (Assuming 20% of the 16b transistors in M1 is CPU part, it would be 3-ish billion transistors, and the Intel does not seem to publish transistor counts but I believe it's more than that for the Intel i9 chips in last year's macbookspro) Perhaps my estimates are wrong, but it seems that they aren't throwing more hardware, but managing to achieve much more with the same "amount of hardware" because it is substantially different.
- temac 6y agoWell you don't have any of the AVX512 nonsense, but probably the OOO of the M1 uses more transistors than on an Intel chip. And Google tells me that a quad core i7-7700K has 2.16 B.
- throwarchitect 6y ago> the M1 is fast isn’t due to technical tricks, but due to Apple throwing a lot of hardware at the problem. Apple threw more hardware at the problem and they lowered the frequency. By lowering the frequency relative to AMD/Intel parts, they get two great advantages. 1) they use significantly less power and 2) they can do more work per cycle, making use of all of that extra hardware.
- baybal2 6y ago> But Apple has a crazy 8 decoders. Not only that but the ROB is something like 3x larger. You can basically hold 3x as many instructions. No other mainstream chip maker has that many decoders in their CPUs. The author completely misses "the baby in the water" Yes, X86 core are HUGE, the whole of CPU is for them only. They can afford wider decode, even though at a giant area cost (which itself would be dwarfed area cost of cache system area.) The thing is, have more decode, and buffer will still not improve X86 perf by much Modern X86 has good internal register, and pipeline utilisation, it's simply they don't have something to keep all of those registers busy most of the time! What it lacks is memory, and cache I/O. All X86 chips today are I/O starved at every level. And that starvation also comes as a result of decades old X86 idiosyncrasies about how I/O should be done.
- coliveira 6y agoHow does I/O works differently in a M1 chip compared to x86?
- baybal2 6y agoI think X86 is the only modern ISA family that still have a separate address space for I/O. It is not used today anymore, but it exists somewhere deep in the chip, and its legacy kind of messed up how the entire wider memory, and cache systems on X86 were designed. X86 has got memory mapped I/O for modern hardware, but on the way there, X86 memory access got tangled with bus access. X86 still treats the wider memory system as a kind of "peripheral" with mind of its own. The intricacies how X86 memory access evolved to keep accommodating decades old drivers, and hardware apparently made a grand mess of what you can, and what you cannot memory map, or cache, and many things deeper in the chip. One of may casualties of that design decision is the X86 cache miss penalty, and an overall expensive memory operations.
- remexre 6y agoWhy can't they just give the IO bus a slower clock and devote the resources to the memory bus? Or, memory-map everything and make the IO area yet another reserved area the BIOS tells the OS about?
- rayiner 6y agoThe answer of wide decode and deep reorder buffer gets much closer than the “tricks” mentioned in tweets. That still doesn’t explain how Apple built an 8-wide CPU with such deep OOO that operates on 10-15 watts. The limit that keeps you from arbitrarily scaling up these numbers isn’t transistor count. It’s delay—how long it takes for complex circuits to settle, which drives the top clock speed. And it’s also power usage. The timing delay of many circuits inside a CPU scare super-linearly with things like decode width. For example, the delay in the decode stage itself scales quadratically with the width of the decoder: ftp://ftp.cs.wisc.edu/sohi/trs/complexity.1328.pdf (p. 15). The delay of the issue queues is quadratic both in the issue width and the depth of the queues. The delay of a full bypass network is quadratic in execution width. Decoding N instructions at a time also requires a register renaming unit that can perform register renaming for that many instructions per cycle, and the register file must have enough ports to be able to feed 2-3 operands to N different instructions per cycle. Additionally, big, multi-ported register files, deep and wide issue queues, and big reorder buffers also tend to be extremely power hungry. On the flip side, the conventional wisdom is that most code doesn’t have enough inherent parallelism to take advantage of an 8-wide machine: https://www.realworldtech.com/shrinking-cpu/2/ https://www.realworldtech.com/shrinking-cpu/2/ (“The first sign that the party was over was diminishing returns from wider and wider superscalar designs. As CPUs went from being capable of executing 1, to 2, to 4, to even 6 instructions per cycle, the percentage of cycles during which they actually hit their full potential was dropping rapidly as both a function of increasing width and increasing clock rate.”). At the very least, such designs tend to be very application-dependent. Branch-y integer code like compilers tend to perform poorly on such wide and slow designs. The M1 by contrast manages to come close to Zen 3, which is already a high ILP CPU to begin with, despite a large clock speed deficit (3.2 ghz versus 5 ghz). And the performance seems to be robust—doing well on everything from compilation to scientific kernels. That’s really phenomenal and blows a lot of the conventional wisdom out of the water. An insane amount of good engineering went into this CPU.
- temac 6y ago> An insane amount of good engineering went into this CPU. I agree but lets not overblow the difficulties either. > For example, the delay in the decode stage itself scales quadratically with the width of the decoder. That could be irrelevant for small enough numbers, and ARM is easier to decode than x86. So this can be very well be dominated by other things. What you cite seems to be only about decoding logical register decoding going into the renaming structures, and then for just that tiny part it even tells that "We found that, at least for the design space and technologies we explored, the quadratic component is very small relative to the other components. Hence, the delay of the decoder is linearly dependent on the issue width." > The delay of a full bypass network is quadratic in execution width. Maybe if that's a problem don't do a full bypass network. > dropping rapidly as both a function of increasing width and increasing clock rate Good thing that the clock rate is not too high then :p More seriously the M1 can keep the beast fed probably because everything is dimensioned correctly, (and yes also because the clocks are not too high, but if you manage to make a wide and slow CPU that actually works well, I don't see why you would want to scale the freq too much high, given you would quickly consume like crazy, and there is only only limited headroom above 3.2GHz anyway). It obviously helps to have a gigantic OOO. So I don't really see where there is so much surprise. Esp. since we saw the progression in the A series. To finish probably TSMC 5nm does not hurt. The competitors are on bigger nodes and have smaller structures. Coincidence? Or just like it has worked during decades already.
- skavi 6y ago> Unlike what has been said on Twitter the answer to why the M1 is fast isn’t due to technical tricks, but due to Apple throwing a lot of hardware at the problem. Of course Apple’s advantage is not solely due to technical tricks, but neither is it entirely, or even mostly due to an area advantage. If it were so easy, Samsung’s M1 would have been a good core.
- davrosthedalek 6y agoHow does 8 wide decode on ARM RISC compare to 4 wide decode on x64 CISC? If, say, you'd need two RISC ops per CISC op on average, that should be the same, right?
- gpderetta 6y agoMost instructions map 1:1. x86 instructions can encode memory operands potentially doubling the practical width but a) not all CPUs can decode them at full width and b) in practice compilers generate mem+ops instructions only for a (not insignificant ) minority of instructions. So the apple to apple (pun intended) practical width difference is closer than it appear, still not as wide. X86 machines usually target 5-6 wide rename, so that would become a bottleneck (not all instructions require rename of course). I expect that M1 has an 8-wide rename. Edit: another limitation is that most x86 decoders can only decode 16 bytes at a time and many instructions can be very long, further limiting actual decode throughput. On the converse the expectation is that most hot code will skip decode completely and is fed from the uop cache. This also saves power.
- GeekyBear 6y agoAnother detail that came out today is just how beefy Apple's "little" cores are. >The performance showcased here roughly matches a 2.2GHz Cortex-A76 which is essentially 4x faster than the performance of any other mobile SoC today which relies on Cortex-A55 cores, all while using roughly the same amount of system power and having 3x the power efficiency. https://www.anandtech.com/show/16192/the-iphone-12-review/2 https://www.anandtech.com/show/16192/the-iphone-12-review/2
- fomine3 6y agoIncredible. Now it should be MiDdLe core.
- Sparkyte 6y agoThank you for a better synopsis. Bugs me so much when people don't look at the logical side of things. Tons of mac'n'knights going around downvoting and stating people are wrong that it has something to do with being a RISC processor. While fundamentally pre 2000s things were more RISC this and CISC that the designs are more similar than ever on x86 and ARM. Just that the components are designed different to handle the different base instruction sets. Also the article is entirely wrong about SoC Ryzen chips have been SoC since their initial inception. In fact SoC on the First APU. Those carried North Bridge components onto the CPU die.
- mixmastamyk 6y agoThe 5nm process is a large factor as well.
- gabereiser 6y agoSpot on. Exactly this. It’s like pre-iPhone when people just assumed you had a laptop and a cellphone. Then Apple said “phone computer!” and changed the game. Same with iPad just less innovation shock. Meanwhile we continued to have this delineation of computer / phone while under the hood - to a hardware engineer - it’s all the same. Naturally the chips they produced for iOS-land are beasts. My phone is faster than the computer I had 5 years ago. My M1 air is just a freak of nature. On par with high end machines but passively cooled and cheaper. I’m still kinda in awe. Not a big fan of the hush hush on Apple Silicon causing us all to play catch-up for support, but that’s Apple’s track record I guess. The M1 is all the things they learned from the A1-A12 chips (or whatever the ordering) which is over a decade of tweaking the design for efficiency (phone) while giving it power (iPad).
- Taniwha 6y agoNot mentioned in the article is the downside of having really wide decoders (and why they're not likely to get much larger) - essentially the big issue in all modern CPUs is branch prediction because the cost of misprediction on a big CPU is so high - there's a rule-of-thumb that in real-world instruction streams there's a branch every 5 instructions or so ... that means that if you're decoding 8 instructions each bundle has 1 or 2 branches in it, if any are predicted taken you have to throw away the subsequent instructions - if you're decoding 16 instructions you've got 3 or 4 branches to predict, chances of having to throw something away gets higher as you go .... there's a law of diminishing returns that kicks in, and in fact has probably kicked in at 8
- legulere 6y agoIt also has shared 12MB of L2 cache for the performance cores which is huge.
- cbsmith 6y ago...and Apple was able to throw hardware at the problem because they got TSMC's manufacturing process. When everyone else is using 5nm, let's see if any of this other stuff actually matters.
- raxxorrax 6y agoThe problem is that a phone with a lightning fast CPU is rather useless with current ecosystems. I do think there are "technical tricks" though, especially the compatibility memory mode that makes x86 emulation faster than comparable ARM chips. If you call it a trick or finesse is probably a matter of perspective.
- Guthur 6y agoM1 is also the only TSMC 5nm chip that is widely available and there is nothing remotely comparable from a process stand point.