5 ms·
“New Intel chips were often delayed and offered only small improvements over previous generations.” As noted on HN and elsewhere, Intel’s struggle has been in
by BooneJS 6y ago
“New Intel chips were often delayed and offered only small improvements over previous generations.”
As noted on HN and elsewhere, Intel’s struggle has been in manufacturing. They minimized risk (“got lazy”) when Moore’s Law reigned and their fabs were 1-2 generations ahead of everyone else, but have since squeezed every last architectural drop out of 14nm. Put Tiger Lake on TSMC 5nm and a lot of M1’s lead goes away.
The M1 is a damned fine chip and I’d like to own one. But the recent hero worship is a bit awkward to read as Apple is generally following an accepted playbook in the End Times of Moore’s law.
1. Save power by creating accelerated logic blocks without locking yourself out of algorithmic improvements.
2. Execute flawlessly. A chip respin costs 3-4 months and who knows how much money, so you need to go to production with the A0 silicon you powered up.
3. Know your software-only workloads. At best, all mathy code can be shoved down SIMD pipelines for high IPC. At worst, you’re dealing with a bunch of branchy integer code. The M1 has solved for both with vector extensions and a massive reorder buffer and register file.
It’s a great chip and Apple got there first. They’ll be rewarded with sales and increased market share, and developers will have one architecture to wrench on for all platforms.
You would be right to say that Intel couldn’t make the M1. It’s not because Intel’s fabs are struggling or Apple has found tricks no one else in the world knows about. It’s because Apple controls the entire software stack and therefore knows what they need on the chip. Intel caters to the entire market (one the M1 just noticeably shrunk) with myriad integrators, workloads, and users, and each has different concerns and priorities making accelerated algorithm blocks useless to some and not-fast-enough for others.
- zepto 6y ago“ You would be right to say that Intel couldn’t make the M1. It’s not because Intel’s fabs are struggling or Apple has found tricks no one else in the world knows about. It’s because Apple controls the entire software stack and therefore knows what they need on the chip. Intel caters to the entire market (one the M1 just noticeably shrunk) with myriad integrators, workloads, and users, and each has different concerns and priorities making accelerated algorithm blocks useless to some and not-fast-enough for others.” This is an important point. It applies up and down the stack for everything Apple does. The people who suggest breaking up Apple are correct that it would destroy them, by forcing them to use inefficient generic components at every level. If we want competition against Apple, the way to get there is not by hobbling them with the inefficiency of the old paradigm. It’s by the rest of the industry figuring out how to collaborate on delivering similar gains without vertical integration.
- alwillis 6y agoIt’s by the rest of the industry figuring out how to collaborate on delivering similar gains without vertical integration. This is one of the smartest things I've seen on HN in a long time. This is exactly what has to happen if the other companies don't want Apple to eat their lunch.
- singhrac 6y agoThe thing is... doesn't Samsung already do this with their Exynos chips? And even with Qualcomm, they must already build them with Android (even a particular phone) in mind, right? Yet somehow the A14 is faster than a Snapdragon 865. I think strategically if Exynos put pressure on Qualcomm in a nontrivial way (e.g. Samsung switched entirely, and started selling Exynos chips to their competitors) that would change the game. Maybe the fab needs to be spun out of Samsung, just as Intel's fabs might be spun out as well.
- zepto 6y agoI don’t see Samsung doing this. It’s one thing knowing that you are going to run Android. It’s entirely another to decide to be able to invest years in optimizing reference counting in collaboration with the compiler team committing to ARC.
- ksec 6y agoSamsung and Qualcomm have different priorities. Vendors are already complaining to Qualcomm about expensive packages due to 5G. They want it cheaper, at the same time people are suggesting Qualcomm's CPU goes faster. You can only have one of them.
- saurik 6y agoThe places where people want Apple broken up are generally places where we all know it will make the result "worse" but will enable interoperability and competition and break down lock-in and empower users and lead to more repairable and less wasteful devices and (and and)... the idea is that what is best for an entire market isn't merely "the fastest thinnest devices with the longest battery life and the most integrated stacks" if that implies buying into one of a handful of locked down systems, but instead that we should actively enforce an optimization that is harmful for all of those companies and potentially their products but which causes an ecosystem of interoperable components where users have commensurate power to make demands back on platforms. It is the very true fact that economies of scale and efficiencies of vertical integration exist that make regulation and anti-trust law so critically important to prevent the world from being made up of an oligopoly of maybe three giant companies that each make one of everything you might ever own, all beautifully integrated with their own products while being almost entirely incompatible with the products made by the other two, with cross promotions from tying and forced purchases from bundling all making it difficult-to-impossible for a new company to independently introduce anything new without replicating the entire vertical stack of an incumbent.
- danpalmer 6y ago> But the recent hero worship is a bit awkward to read as Apple is generally following an accepted playbook in the End Times of Moore’s law. Is anyone else doing this though? It might be the theoretical playbook, but if others aren't doing this in practice then Apple are worthy of credit for pulling it off. 1 - accelerators for specific things aren't new, and lots of processors have an H264 accelerator, but Apple seems to be pushing this a bit further than others. Even low power vs high power cores doesn't seem to be something others are doing much of. 2 - I don't know enough about. 3 - Apple has put a ton of FP units on this seemingly for JS performance, anyone else could be doing that, it's not a big spoiler to know that people run a lot of JS now, and yet the M1 has more FP units than any other consumer CPU. This knowledge about the software workloads seems to be beyond what others are working on. It seems to be more than just the 5nm process, not least because the A14 generation is the first 5nm chip but Apple's chips have been well respected back to 7 and 10nm for similar reasons. Caveat, very much a layman on this, but it seems like Apple are executing that playbook significantly better than others.
- ksec 6y ago>Is anyone else doing this though? Yes, but from a different perspective. For example the recent M&A between lots of Fabless companies are precisely because of that. Sharing cost and allowing different IPs to work together. What is missing is someone doing the Software. If it was Microsoft, they would have done at least done something. It wouldn't be good, but they would at least be "trying" to compete. Google just doesn't care.
- klelatti 6y agoAgree 100% but some of the hero worship is a reaction I think to being told for a long time that 'Only Intel x86 chips are powerful enough' - which in part was put about by Intel marketing (Intel inside etc and some of the comments at the time the iPhone launched that consumers need an x86 to do real work) - and that ARM wouldn't scale up for desktops etc. The other key point is that Apple can pick and mix from the technologies they use - they've unbundled Intel - choosing the best that they can find for their needs whether in-house or externally.
- majormajor 6y agoWhat workloads are you picturing that Apple isn't serving with Mac chips that Intel has to cater to? This has to be a much more general purpose device than an iPhone.
- r00fus 6y ago> 2. Execute flawlessly. You say this like it's easy or routine. At the bleeding edge, execution is 99% the hard part.
- mikechen233 6y agoI found say intel not just failed in the process nodes. It also failed to see the rise of the dedicated fab (tsmc) and fabless business model as a more efficient model. Intel also failed at innovating on lower power designs and the gpu. Even at 14nm, there are so much that could be done to improve igpu performance. From color compression, to the more efficient shader array. The nintendo switch runs Nvidia tegra x1. It's a 20nm chip, and look at the performance of the gpu. Intel could have also put their markshare weight to put the alternative to cuda be it opencl or even their own API to really gain market traction. They could have their dedicated gpu. Yet they just ignored the whole massive simd compute market. I think for intel, lossing mobile to arm and see the improvements in mobile year over year should be a wake-up call. Yet they didn't do anything to increase their moat for the past 10 years. Or maybe they did try, but their strategies are just flawed. The strategists at Intel just don't have enough foresight and vision.
- klelatti 6y agoFeels like execution rather than strategy: they've tried at mobile, gpus, foundry etc - the right calls - but just haven't delivered. Apple / AMD / TSMC have executed much better.
- mikechen233 6y agoI feel is both. Also a good strategist would have made sure whatever plan and vision you have is also going to be executed well. A strategy that didn't end up to be executed well is a failed strategy.
- Lio 6y agoInterestingly, Intel could also control the whole stack. They make NUC hardware[1] and Intel Clear Linux[2]. They could tune that stack for a new purpose designed chip... if they wanted to, but it seems unlikely. 1. https://www.intel.com/content/www/us/en/products/boards-kits/nuc.html https://www.intel.com/content/www/us/en/products/boards-kits... 2. https://clearlinux.org/ https://clearlinux.org/
- kinghajj 6y agoIt's not just the re-order buffer, the M1 also has 8-way decode (vs 4-way for Intel & AMD microarchitectures), and a 128 KiB L1D cache (vs 48 recently from Intel, and 32 from AMD), and that cache still has a 3-cycle latency. These all let an M1 chip keep up in single-threaded workloads @ 3.2 GHz with what the latest Zen 3 5950X can do @ 5.0 GHz.
- EricE 6y ago"Intel caters to the entire market (one the M1 just noticeably shrunk) with myriad integrators, workloads, and users, and each has different concerns and priorities making accelerated algorithm blocks useless to some and not-fast-enough for others." Which is why even if Intel does manage to solve their process problems, they are not out of the woods yet. As you point out Apple's per core performance is impressive, but their ability to customize their hardware to optimize very specific compute tasks like the nsobject release example or floating for Javascript will be a lot harder to convince off the self/least common denominator parts like x86 processors have traditionally been. I also think we are at the beginning of seeing the influence of "accelerators" such as the ML and neural engine cores. Accelerators beyond the CPU are nothing new - that's what a graphics card is, after all. What's interesting is these are on the same SOC with the CPU cores. What if Apple decides even for the higher end Mac's to integrate RAM into or closely couple with the SOC? What if they decide on their own memory controller that lets you address all the memory chips at once. What potential could that have for performance? Maybe not enough to offset the cost/complexity. But if it did turn out to be on the right side of the cost/benefit equation now instead of having to convince multiple parties it's a good idea, at worst Apple just has to convince a few internal divisions to work together. That is the long term power I see with Apple Silicon. Not raw specs or Moores Law - but unfettered access to the full Art of the Possible.
- SubjectToChange 6y agoAs you point out Apple's per core performance is impressive, but their ability to customize their hardware to optimize very specific compute tasks like the nsobject release example or floating for Javascript will be a lot harder to convince off the self/least common denominator parts like x86 processors have traditionally been. The improvement in retaining and releasing NSObjects is mostly due to Arm's weaker memory model. And javascript specific instructions like FJCVTZS are available for anyone designing Arm cores. Also, IIRC, FJCVTZS was added to more closely match the way x86 handles floating point numbers because that is what Javascript was designed to target. But if it did turn out to be on the right side of the cost/benefit equation now instead of having to convince multiple parties it's a good idea, at worst Apple just has to convince a few internal divisions to work together. The problem is that Apple does what they want at their discretion. You need them, but that don't need you. Whereas hardware manufactures are usually quite sensitive to customer needs. We've already seen how this plays out with Apple's professional software products like Shake or Final Cut. Not to downplay Apple's accomplishment, after all nobody expected them to be this competitive. But people are getting caught up in the hype.
- gumby 6y ago> Put Tiger Lake on TSMC 5nm and a lot of M1’s lead goes away. > It’s not because [...] Apple has found tricks no one else in the world knows about. I don't think anyone would disagree with either statement. The point is Intel can't get TL on a small node, and not for lack of trying. They eventually will, of course, but there is an organizational sickness at Intel that has crippled them.
- mlindner 6y agoThen how is M1 beating _desktop_ top end AMD chips in single core performance? Ignoring Intel here.
- rasz 6y agoIt isnt. In Cinebench its up there behind middle of the road Ryzen 5600X in single core, and ~two times slower in multi.