4 ms·
Hmm. Not very interesting take. First, the M1 is the 12th chip design in the Apple Silicon series (excluding the X variants). They have cranked out a more ca
by dev_tty01 5y ago
Hmm. Not very interesting take. First, the M1 is the 12th chip design in the Apple Silicon series (excluding the X variants). They have cranked out a more capable, higher performing, lower power chip every year since 2010.
https://en.wikipedia.org/wiki/Apple_silicon https://en.wikipedia.org/wiki/Apple_silicon
Second, it's not about clock rate. That is only one small part of the story. It is really about instructions per cycle per core. Apple is killing it on that front and running wider at lower rates is a big part of how they are outperforming in performance per watt while still winning in single core performance. We may see some clock rate increase in an M2, but I suspect their basic design philosophy won't change. It is just working too well.
For the curious, see https://travisdowns.github.io/blog/2019/06/11/speed-limits.html#ooo-table https://travisdowns.github.io/blog/2019/06/11/speed-limits.h....
That table shows that the M1 has a much bigger reorder buffer, large load and store buffers, huge integer and vector register files, way more branches in flight, etc. By eschewing high clock rates, they are able to really go after massive concurrency at the hardware level in a single core. 7 simultaneous integer operations, 4 simultaneous floating point, multiple load and store. It's a beast.
https://www.anandtech.com/show/16226/apple-silicon-m1-a14-deep-dive/2 https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
Of course they will come out with an M2 soon. They've been doing this year after year for over a decade.
- saagarjha 5y ago> First, the M1 is the 12th chip design in the Apple Silicon series How are you counting that?
- BugsJustFindMe 5y agoA4,A5,A6,...,A14,M1
- mastazi 5y agoI'm not parent but the A series starts with A4 which was the first commercially released model. So between A4-A14 we have 11 generations, and M1 is the 12th.
- matthewfcarlson 5y agoTechnically there's all the X versions which are boosted versions and are somewhat different internally.
- saagarjha 5y agoA14 and M1 are from the same series.
- mastazi 5y agoI guess you mean same generation. Apple uses the word "series" to identify A Series, M Series etc. By the way, yes, both A14 and M1 use Firestorm + Icestorm cores: so I guess that they are the same generation from a technical standpoint.
- deleted 5y ago[deleted]
- soperj 5y ago> They've been doing this year after year for over a decade. So was Intel for decades, then they weren't.
- saagarjha 5y agoSure, but we saw them stalling for several years. Apple doesn’t seem like they’re in that sort of rut.
- bredren 5y agoNote, Intel had three different CEOs in the past four years.
- totalZero 5y agoApple was also ballasted by iPhone sales during the lull in personal computing that took place about five years ago. Take a look at their income breakdown and you'll see what I mean. Intel bungled the transition to EUV but that isn't the only challenge they have had to surmount in the past decade. Personally I don't see the Intel rut as particularly deep or mucky. Intel has good management, and Gelsinger has deep knowledge of how enterprise customers operate due to his experience. They have a road map for some exciting product releases in the next couple of years, and they dominate their game in terms of market share. Intel made over $20B profit last year, and semiconductor demand is booming across the board, but they still get trashed by the masses. It's really interesting (if you're into stocks) to compare Intel’s P/E ratio against the rest of the sector. Even the market doesn't think particularly highly of them.
- zucker42 5y agoIntel's stalled mostly at the fabrication side. Since Apple is fabless, this isn't really a good comparison.
- kmonsen 5y agoI don't think this is a story, but someone still has to fab them and they can stall
- gopalv 5y ago> 7 simultaneous integer operations, 4 simultaneous floating poin The floating point is going to be an interesting thing to look at - the CPUs made targeting HPC workloads tend to be flops heavy, but the flops tend to be starved for memory bandwidth unless you're doing exactly the best vector processing you can & that fortran can do a great job with. So you throw in a lot more oomph on the vector side and leave single operation float arithmetic at 2. Floating point operations of a smaller size of values (more realistically, quarternions or rgba) would be the reason that M1 feels a little bit more snappy when it comes to basic things like text-layout code or graphics images which don't do SIMD very well, but still consume a lot of arithmetic. I'd suspect that the vertical integration is going to be the secret, because it looks like more profile information of desktop apps going into chip design here. A similar story is expected of the Graviton series as well, with AWS having a good idea what to build for.
- zelon88 5y agoWhy doesn't anyone ever bother to mention that the M1 is a RISC CPU when discussing that it can do more IPS than x86? There are 1024 possible Armv8 instructions [1] as opposed to 1,503 x86 instructions [2] and 3,684 x86-64 instructions [3]. There are things x86 and x86-64 can do in a single instruction that would take dozens of instructions to accomplish on Arm. [1] https://www.csie.ntu.edu.tw/~cyy/courses/assembly/10fall/lectures/handouts/lec09_ARMisa.pdf https://www.csie.ntu.edu.tw/~cyy/courses/assembly/10fall/lec... [2] https://fgiesen.wordpress.com/2016/08/25/how-many-x86-instructions-are-there/ https://fgiesen.wordpress.com/2016/08/25/how-many-x86-instru... [3] https://www.csie.ntu.edu.tw/~cyy/courses/assembly/10fall/lectures/handouts/lec09_ARMisa.pdf https://www.csie.ntu.edu.tw/~cyy/courses/assembly/10fall/lec...
- wmf 5y agoIn real workloads ARM executes only slightly more instructions than x86.
- sudosysgen 5y agoWhile true, it does also mean more stress on the reordering system, cache, branches, etc... Than on x86.
- SilverRed 5y agoDo number of instructions even matter on modern chips? intel/amd chips don't even use x86 since they break it down in to their own internal RISC via a translation layer? If you have a CPU that breaks one mega x86 instruction in to 100 internal instructions, is that any better than 100 external instructions generated by a compiler?
- sudosysgen 5y agoIt does and doesn't. When your CPU is breaking up that instruction you can optimize for the exact way its being broken up at the design because you know what those 100 instructions will be and in what order. You pay for that with more decode. It's a trade-off, but it does mean there is less of a need for the machinery the M1 has more of, somewhat.
- sudosysgen 5y agoThe M1 can indeed run more instructions at once and branch those instructions better, but it comes at a cost. That cost is that it's larger. At the same power AMD can stuff twice the cores, with similar though lower single core performance. That also means lower clock speeds, and there are less instruction level guarantees you can rely on thus somewhat more complexity is necessary. And no, it's not really a beast. It's competitive.
- leucineleprec0n 5y ago>At the same power AMD can stuff twice the cores Note that the “TDP” is meaningless for actual performance/watt comparisons with a controlled parameter due to the variation in this term. It’s just a marketing term. The Ryzens on 7NM consume much more power than Apple’s Big cores did on the A12 and A13, both of which were on TSMC 7NM. They are also both competitive with the single-core scores of the Ryzen Zen 2 or 3 cores, if not equivalent while being older architectures than AMD’s. The M1’s Big [Firestorm] cores also consume less power and achieve more performance than the Zen 3 core. It’s safe to say Apple’s architecture is largely superior, outright. https://images.anandtech.com/doci/16226/spec2006_A14.png https://images.anandtech.com/doci/16226/spec2006_A14.png
- sudosysgen 5y agoI'm not talking about TDP when I talking about stuffing cores, I mean physical size. The M1 cores use more transistors. I'm sorry, but you're comparing power draw of a workstation "X" AMD chip with a laptop chip. It's simply not a honest comparison. You must compare the efficiency of mobile chip with a mobile chip. When you do that you find similar efficiency. I don't understand why no one is posting actual apples to apples comparison. Every time a comparison is posted its either comparing to a workstation chip to find power efficiency even if they're tuned to be power inefficient, or comparing to Intel CPUs only, etc... Laptop processors from AMD use less than half the power per core of workstation processors while sacrificing only a very small performance gain, due to the use of a different lithography for I/O. Apple's architecture is simply not superior. If it was, we wouldn't be making these contrived comparisons, and we wouldn't even be comparing 5nm chips to 7nm chips.
- saati 5y ago> That table shows that the M1 has a much bigger reorder buffer, large load and store buffers, huge integer and vector register files, way more branches in flight, etc. That leaves out the most important thing that enables all of that, the 8-wide symmetric decoder that can feed those. x86_64 cpus only have 5-wide ones, and only the first of those can decode the multi-uops instructions, and even worse there are even more complex instructions that are microcoded.
- dev_tty01 5y agoYes, absolutely right. That decoder is also beastly. Of course, fixed (nearly) length ARM instructions make the problem a lot easier. x86 instruction format and length variance is just ridiculous. Intel has saddled themselves with years of complexity and it has become a big obstacle for them. They have always muscled through it by throwing process innovation at the problem, but recent stumbles leave them trapped in a box of their own creation. Intel is a great company and I expect them to dig out of this, but it is tough right now.