9 ms·
Put another way: if AMD (and especially Intel) don't do something about this they're going to get completely eaten alive by ARM. The amount of processing power
by cletus 4y ago
Put another way: if AMD (and especially Intel) don't do something about this they're going to get completely eaten alive by ARM.
The amount of processing power available in a modern smartphone is truly mind-boggling. I'd love to see a chart showing the chip cost and energy cost of the power on an M1 chip in each previou syear. I would guess that 30+ years ago you'd be in the millions of dollars and watts of power but that's just a guess.
As we see from the modern M1/M2 Macbooks, these lower TDP SoCs are more than capable of running a computer for most people for most things. The need for an Intel or AMD CPU is shrinking. It's still there and very real but the waters are rising.
- beebmam 4y agoIn my own experience, the supposed ARM chip superiority claims are almost entirely marketing. I get significantly better performance (15-50%) from nearly all of my CPU workloads on modern Intel/AMD hardware vs the ARM Apple devices.
- CalChris 4y agoThe article is about energy efficiency. Do you get 15-50% better performance per watt from nearly all of your workloads?
- Salgat 4y agoIf you take a Zen 3 running at optimal clocks for efficiency (such as the 5800U) the difference between its computing performance per watt is competitive with the M1 if you account for the difference in node size (which TSMC claims gives 30% less power consumption at the same performance). As the article points out, the real efficiency gains will be domain specific changes such as shifting to 8 bit for more calculations.
- jorvi 4y agoHow would x86-64 be as efficient with the same transistor & power budget when they have to run an extra decoder and ring within that budget? Seems physically impossible.
- Salgat 4y agoMy guess is it's related to the higher transistor count. The M1 for example has 16B transistors compared to the 5800U with 10.7B.
- Panzer04 4y agoAs I understand it, the actual processing part of most chips nowadays is fairly bespoke, with a decoder sitting on top. I doubt decode can make up that large a portion of a chips power consumption (probably negligible next to the rest of the chip?), so other improvements can make up for the difference.
- ejiblabahaba 4y agoI found this to be a pretty expansive answer to this question: https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-matter/ https://chipsandcheese.com/2021/07/13/arm-or-x86-isa-doesnt-...
- jorvi 4y agoThank you for the detailed article!
- jabl 4y agoAll else being equal, they can't. But the difference isn't as big as some people like to think. For a current high end core, probably low single digit %. And x86-64 has had a lot more effort going into software optimization.
- monocasa 4y agoBecause the more complex decoder is traded in this case for a denser instruction set, which means they can trade it for less instruction cache (which is more power hungry).
- chasil 4y agoHow can AArch64 be as efficient when it implements all of the old 32-bit extensions? Some don't, but a phone does.
- senttoschool 4y agoIt really isn't competitive. First of all, when you downclock anything, you're going to gain efficiency. If Apple downclocks M1, it can get even more efficient. Second, most of these tests use Cinebench, which is highly optimized for x86, not ARM instructions. Geekbench should be used instead. Third, the M1 is a SoC. Everything is on it. Everything is efficiently connected directly inside the chip.
- Salgat 4y agoBoth the M1 and 5800U run around 15W and are already clocked for efficiency. The M1 Max is their higher clocked less efficient offering.
- senttoschool 4y agoThis is false. The 5800U will boost well beyond 15w. Ignore their TDP marketing ratings.
- cookiengineer 4y agoHonestly I don't understand why there's not something like a 256 core ARM laptop with 4TB RAM. The benefit of ARM is scale of multitasking due to not requiring the same kind of lock states that Intel's architecture requires, and can additionally scale much better than only one physical+virtual core pair. I guess the only thing that's holding back ARM is Microsoft, as laptops are expected to run an desktop OS that people are comfortable with. Windows RT wasn't really a serious desktop OS and rather a joke made only for some IoT enterprises instead of end-users. I wish there was more serious hardware than the standard broadcom or MediaTek chips, I'd definitely want some of that...be it as a mini ATX desktop/server format (e.g. as a competitor to Intel NUC or Mac Mini) or as a laptop. With the ongoing energy crisis something like solar powered servers would be so much more feasible than with x86 hardware.
- PragmaticPulp 4y ago> Honestly I don't understand why there's not something like a 256 core ARM laptop The high power ARM cores aren’t that small. If you took the M2 and scaled it up to 256 cores, it would be almost 7 square inches. You can’t just scale a chip like that, though, so the interconnects would consume a huge amount of space as well. It would also consume over 1000W. The latest ARM chips are great, but some times I think the perception has shifted too far past the reality.
- deleted 4y ago[deleted]
- Dylan16807 4y ago7 square inches would also include an enormous GPU and tons of accessories. The actual cores are about .6/2.3 mm², and local interconnects and L2 roughly double that. So with just those parts, 256 P-cores would be about 1.5 square inches, and 256 E-cores would be about half a square inch. And in practical terms you can fabricate a die that's a bit more than a square inch. Of course it wouldn't use 1000 watts. When you light up that many cores at once you use them at lower power. And I doubt a 256 core design would have all that many P cores either. As a rough estimate, you could take the 120mm² M1 chip, add 28 more P-cores with 110mm², 220 more E-cores with 300mm², 128 more MB of L3 cache with 60mm², 100mm² of miscellaneous interconnects, and still be on par with a high end GPU. That sounds doable but is pushing it. A 128 core die, though, has nothing stopping it except market fit.
- BlueTankEngine 4y agoIn my experience talking to semiconductors folks, ARM is just not a concern anymore. The future is RISC-V, and ARM is already being seen as legacy tech. ARM's progress in the server space has stalled, the ARM Windows ecosystem is dead, Android has laid the groundwork for a move to RISC-V, and ARM has never and will never touch the desktop market.
- pnpnp 4y ago> ARM has never and will never touch the desktop market That’s a bold statement as I type all day on an M1 Mac. My FT100 company just made the leap to them as dev machines.
- r00fus 4y agoMy company has entire teams and regions that do NOT buy PCs and only Mac laptop for employees. Started with the M-series.
- BlueTankEngine 4y ago[flagged]
- neltnerb 4y agoAs an experiment quite a few years ago I got a laptop with a special version of the Intel CPU that was not as fast but much more power efficient. ASUS UL30A-X5 Really an excellent computer, ran linux great (games didn't really exist yet though), and with tuning was coming in under 10W if the display brightness was turned down. First time I was able to get through flights without the system running dead. I think in this case what's going on is that temperature rises increase resistance in a chip and therefore cause lower efficiency. If you can keep it cool, you can keep it more efficient. The move seems like a necessary one, a computer as powerful as that UL30A is probably inside the phone if you turn off the radio and display, that thing still had a giant battery and only lasted 10-12 hours. I've seen AMD do some pretty impressive things, I wouldn't count them out. They're at least willing to attempt to compete on price.
- PragmaticPulp 4y ago> Put another way: if AMD (and especially Intel) don't do something about this they're going to get completely eaten alive by ARM. AMD’s latest parts are actually quite close to M1/M2 in computing efficiency when clocked down to more conservative power targets. They crank the power consumption of their desktop CPUs deep into the diminishing returns region because benchmarks sell desktop chips. You can go into the BIOS and set a considerably lower TDP limit and barely lose much performance. Where they struggle is in idle power. The chiplet design has been great for yields but it consumes a lot of baseline power at idle. M1/M2 have extremely efficient integration and can idle at negligible power levels, which is great for laptop battery life.
- brigade 4y agoPeople keep repeating that Zen4 and M1 are close in efficiency but what is the source with actual benchmarks and power measurements? At any rate, using single points to compare energy efficiency isn't a good comparison, unless either the performance or power consumption of the data points comparable. Like, the M1's little cores are 3-5x even more efficient when operating in an incomparable power class, and Apple's own marketing graphs show the M1's max efficiency is also well below its max performance [1] Those perf/power curves are the basis of actually useful comparisons; has anyone plotted some outside of marketing materials? It might even be possible under Asahi. [1] https://www.apple.com/newsroom/2021/10/introducing-m1-pro-and-m1-max-the-most-powerful-chips-apple-has-ever-built/ https://www.apple.com/newsroom/2021/10/introducing-m1-pro-an...
- Panzer04 4y agoGenerally you see this in the lower class chips that aren’t overclocked to within an inch of instability. It’s not uncommon to see a chip that uses 200w to perform 10% worse at 100w, or 20% worse at 70w. I can’t be bothered to chase down an actual comparison, but usually you’ll see something along those lines if you compare the benchmarks for the top tier chip with a slightly lower tier 65w equivalent.
- scns 4y ago
- phkahler 4y ago>> I would guess that 30+ years ago you'd be in the millions of dollars and watts of power but that's just a guess. 30 Years ago I don't think the compute power of a modern phone chip was available at any price, even in super computers. On a tangential note, there are economists who think this increase in compute is somehow an increase in one of their measures - I don't recall which one. I disagree, because with that logic we all have trillion dollar tech in our pocket. Making a better product over time is expected, it's not some kind of increase in output.
- jodrellblank 4y agoThe Top500 supercomputer list started in June 1993, just about 30 years ago. At the top is the CM-5/1024 by Thinking Machines Corporation at Los Alamos National Laboratory with 1,024 cores and peaking at 131.00 GFlop/s (billion floating point operations per second). It's an Apples to ThinkingMachine Oranges comparison but CPU-Benchmark[1] ranks the Apple A16 Bionic used in the latest iPhones, its GPU - in the "iGPU - FP32 Performance (Single-precision GFLOPS)" section - at 2000 GFlop/s. GadgetVersus[3] reports a GeekBench score of the A16 Bionic at 279.8 GFlop/s. - SGEMM test of matrix multiplication, it seems. AnandTech[4] was reporting the A15 architecture ARMv7 came in at 6.1 GFlops in the "GeekBench 3 - Floating Point Performance" table, SGEMM MT test result, in 2015. [1] https://www.top500.org/lists/top500/1993/06/ https://www.top500.org/lists/top500/1993/06/ [2] https://cpu-benchmark.org/cpu/apple-a16-bionic/ https://cpu-benchmark.org/cpu/apple-a16-bionic/ [3] https://gadgetversus.com/processor/apple-a15-bionic-gflops-performance/ https://gadgetversus.com/processor/apple-a15-bionic-gflops-p... [4] https://www.anandtech.com/show/8718/the-samsung-galaxy-note-4-exynos-review/6 https://www.anandtech.com/show/8718/the-samsung-galaxy-note-...
- phkahler 4y agoInteresting. I would have thought a few GFLOPs today would have been faster than the old super computer, but nope. The GPU is faster though. Still, the phone has both and can run on battery power while fitting in your pocket ;-)
- tails4e 4y agoFrom what I've seen the CPU/ALU/decode at the center being ARM or x86 may make less difference than you think. The amount of circuitry and silicon area (correlated with power) for non core is significant. MMUs, vector instructions, complex cache hierarchy, high speed IO (DDR, pcie, you name it) extremely complex network on chip (infinity fabric) to enable cpu interconnectivity, etc. Is very significant. Look at the IO die size vs the CCD size. As one poster pointed out using chiplets have great advantages, but there is a power hit. Thankfully newer tech is bringing that power down too. I'd love to see a power breakdown of a full chip to see what % is attributed to the cpu core itself.
- deleted 4y ago[deleted]