33 ms·
Single-chip processors have reached their limits
- AnimalMuppet 5y ago"More multi-chip processor designs" != "single-chip processors have reached their limits".
- gnarbarian 5y agoAre we moving this way because bigger chips with many cores have worse yields? so the answer is to make lots of little chips and fuse then together?
- deleted 5y ago[deleted]
- monocasa 5y agoYeah, in the very general case, chip errors are a function of die area. Cutting a die into four pieces so that when an error occurs in manufacturing, you only throw out a quarter of the die area is becoming the right model for a lot of designs. Like all things chips, it's way more complicated than that, fractally, as you start digging in. Like AMD started down this road initially because of their contractual agreements with GloFlo to keep shipping with GloFlo does, but wanted the bulk of the logic on a smaller node than GloFo could provide, hence the IO die and compute chiplets model that still exists in Zen. It's still a good idea for other reasons but they lucked out a bit by being forced in that direction before other major fabless companies. This is also not a new idea, but sort of ebbs and flows with the economics of the chip market. See the VAX 9000 multi chip modules for an 80s take on the same ideas and economic pressures.
- WithinReason 5y agoTheir GPUs are likely to be multichip for the first time too with NAVI 31 (while Nvidia's next gen will still be single chip and likely fall behind AMD). It also seems like that the cache will be 6nm while the logic will be 5nm and bonded together with some new TSMC technology. At least that can be inferred from some leaks: https://www.tweaktown.com/news/84418/amd-rdna-3-gpu-engineer-confirms-hybrid-5nm-6nm-nodes-on-navi-3x/index.html https://www.tweaktown.com/news/84418/amd-rdna-3-gpu-engineer...
- monocasa 5y agoThere's a few ways to interpret that. Another interpretation could be that they are simply taping out Navi32 on two nodes, perhaps for AMD to better utilize the 5nm slots they have access to. Perhaps when Nvidia is on Samsung 10nm+++, then the large consumer AMD GPUs get a node advantage already being at TSMC 7nm+++, and so they're only using 5nm slots for places like integrated GPUs and data center parts that care about perf/watt. But your interpretation is equally valid with the information we have AFAICT.
- ceeplusplus 5y agoI've yet to see any sort of research out of AMD on MCM mitigations for things like cache coherency and NUMA. Nvidia on the other hand has published papers as far back as 2017 on the subject. On top of that even the M1 Ultra has some rough scaling spots in certain workloads and Apple is by far ahead of everyone else on the chiplet curve (if you don't believe me, try testing lock-free atomic load/store latency across CCX's in Zen3). Also AMD claimed the MI250X is "multichip" but it presents itself as 2 GPUs to the OS and the interconnect is worse than NVLink.
- thissiteb1lows 5y ago
- IshKebab 5y agoThis is what Tesla's Dojo does (it's really a TSMC technology that they are the first to utilize). You can cut your wafer up into chips, ditch the bad ones, then reassemble them into a bigger wafery chip thing using some kind of glue. Then you can do more layers to wire them up. I think they do it using identical chips but I guess there's no real reason you couldn't have different chips connected in one wafer. Expensive though!
- sliken 5y agoWell fuse is one possibility. The AMD Epyc has generally an IO+memory controller die (called IOD) + 8 chiplets that are 8 cores each for most of the Epyc chips, however not all cores are enabled depending on the SKU. However apple's approach does allow impressive bandwidth, 2.5TB/sec which is much higher than any of the chiplet approaches I'm aware of.
- kzrdude 5y agoIt's surprising it took that many cores before the limit was reached!
- retrac 5y agoThe best chiplet interconnect may turn out to be no interconnect at all. Wafer scale integration [1] has come up periodically over the years. In short, just make a physically larger integrated circuit, potentially as large as the entire wafer -- like a foot across. As I understand it, there's no particular technical hurdle, and indeed the progress with self-healing and self-testing designs with redundancy to improve yield for small processors, also makes really large designs more feasible than in the past. The economics never worked out in the favour of this approach before, but now we're at the scaling limit maybe that will change. At least one company is pursuing this at the very high end. The Cerebras WFE-2 [2] ("wafer scale engine") has 2.6 trillion transistors with 800,000 cores and 48 gigabytes of RAM, on a single, giant, integrated circuit (shown in the linked article). I'm just an interested follower of the field, no expert, so what do I know. But I think that we may see a shift in that direction eventually. Everything on-die with a really big die. System on a chip, but for the high end, not just tiny microcontrollers. [1] https://en.wikipedia.org/wiki/Wafer-scale_integration https://en.wikipedia.org/wiki/Wafer-scale_integration [2] https://www.zdnet.com/article/cerebras-continues-absolute-domination-of-high-end-compute-it-says-with-worlds-hugest-chip-two-dot-oh/ https://www.zdnet.com/article/cerebras-continues-absolute-do...
- nynx 5y agoThere are some new chip manufacturing technique coming down the pipeline, which will lead to prices dropping and likely "wafer-scale" will get to the mainstream.
- tragictrash 5y agoCould you elaborate? Would love to know more.
- nynx 5y agoUnfortunately, I cannot.
- AbbeSomething 5y agoCan you clarify what you mean? How can anything that would use a major share of a wafer or even a whole one be mainstream? Producing a wafer is expensive, the chips are only available to ordinary people in high income markets because you get hundreds or even thousands from a single wafer.
- tempnow987 5y ago"Reached their limits" - I feel like I've heard this many many times before. Not that I doubt it, but just I've also been impressed with the ingenuity that folks come up with in this space.
- aeturnum 5y agoI read articles like this as saying "reached their limits [as we currently understand them]." Sometimes we learn we were mistaken and more is possible but it's not reliable and, crucially, when it happens it happens in unexpected ways. The process of talking about when (and why) techniques have hit their useful limits is often key to unearthing the next step.
- tawaypol 5y ago"There's plenty of room at the bottom."
- marcosdumay 5y agoActually, we are getting out of room there. that speech is about 80 years old nowadays. There was plenty of room at that time. Of course, it also speculated that we would move into quantum computers at some point, what is still a possibility, but now we know that quantum computers won't solve every issue.
- mjreacher 5y agoAgreed. I would be wary of reaching fundamental limits set by physics although I don't think we're there yet. "It would appear that we have reached the limits of what is possible to achieve with computer technology, although one should be careful with such statements, as they tend to sound pretty silly in five years." - attributed to von Neumann, 1949.
- gameswithgo 5y ago
- syntheweave 5y agoWe only have to solve one limitation per year to keep making progress year over year, and as it is, the semiconductor industry still seems to be solving large numbers of significant issues yearly. So while we don't necessarily get smooth, predictable improvement, a safe bet is that there will be continue to be useful new developments 10-20 years out, even if they don't translate to the same kinds of gains as in years past.
- RcouF1uZ4gsC 5y ago> UCIe is a start, but the standard’s future remains to be seen. “The founding members of initial UCIe promoters represent an impressive list of contributors across a broad range of technology design and manufacturing areas, including the HPC ecosystem,” said Nossokoff, “but a number of major organizations have not as yet joined, including Apple, AWS, Broadcom, IBM, NVIDIA, other silicon foundries, and memory vendors.” The fact that the standard doesn’t include anyone who is actually building chips makes me very pessimistic about it.
- ranger207 5y agoLooks like a lot of people who actually build chips are in the organization https://www.uciexpress.org/membership https://www.uciexpress.org/membership
- Veliladon 5y agoThe M1 Ultra is fabricated as a single chip. The 12900K is fabricated as a single chip and is still a quarter the size of the M1 Ultra. Ryzen 3 puts 8 cores on a CCX instead of four because DDR memory controllers don't have infinite memory bandwidth (contrary to AMD's wishful nomenclature) and make shitty interconnects between banks of L3. Chiplets are valid strategies that are going to be used in the future but there are still more tricks that CPU makers have up their sleeves that they need to use out of necessity. They're nowhere near their limits.
- 2OEH8eoCRo0 5y ago> The M1 Ultra is fabricated as a single chip. I'm curious how much the M1 Ultra costs. It's such a massive single piece of glass I'd guess it's $1,200+. If that's the case it doesn't make sense to compare the M1 Ultra to $500 CPUs from Intel and AMD.
- mrtksn 5y agoWouldn't the price be primarily based on capital investment and not so much on the unit itself? After all, it's essentially a print out on a crystal using reeeeeally expensive printers. AFAIK Apple's relationship with TSMC is more than a customer relationship.
- 2OEH8eoCRo0 5y agoIn a parallel universe where Intel builds and sells this CPU- what's the price? Single chip, die size of 860 square mm, 114 billion transistors, on package memory. It just got me thinking the other day since all of these benchmarks pit it against $500-$1000 CPUs and it doesn't seem to fall in that price range at all. Look at this thing: https://cdn.wccftech.com/wp-content/uploads/2022/03/2022-03-19_14-00-05-2060x1074.png https://cdn.wccftech.com/wp-content/uploads/2022/03/2022-03-...
- mrtksn 5y agoThat's the whole package though, together with the RAM and everything. The Actual die is about the size of the thermal paste stain on that picture.
- refulgentis 5y agoI'm embarrassed to admit I still don't quite understand what a chiplet is, would be very grateful for your input here. If a thread can run on multiple chiplets then this is awesome and seems like a solution. If one thread == one chiplet, then*: - a chiplet is equivalent to a core, except with speedier connections to other cores? - this isn't a solution, we're 15 years into cores and single-threaded performance is still king. If separating work into separate threads was a solution, cores would work more or less just fine.** * put "in my totally uneducated opinion, it seems like..." before each of these, internet doesn't communicate tone well and I'm definitely not trying to pass judgement here, I don't know what I'm talking about! ** generally, for consumer hardware and use cases, i.e. "I am buying a new laptop and I want it to go brrrr", all sorts of caveats there of course
- hesdeadjim 5y agoA chiplet is a full-fledged CPU with many cores on it. The term is used when multiple of these chips are stitched together with a high speed interconnect and plugged into the single socket on your motherboard. If you ripped the lid off a Ryzen "chip", you would see multiple CPU dies underneath for the high end models.
- tenebrisalietum 5y agoAdditionally - MCM - multi-chip module - instead of putting separate chips for various functions on a board, they're fused together in what from the outside looks like a single chip, but internally is 3 or 4 unrelated chips. Examples at the Wikipedia article: https://en.wikipedia.org/wiki/Multi-chip_module https://en.wikipedia.org/wiki/Multi-chip_module
- deleted 5y ago[deleted]
- sliken 5y agoAMD Epyc is (AFAIK) what popularized the term. Their current design has a memory controller (PCIe controller, 8 x 64 bit channels of ram, etc) and 8 chiplets which are pretty much just 8 cores and a infinity fabric connection for a cache coherent connection to other CPUs (in the same or other sockets) and dram. So generally Epyc come with some multiple of 8 CPUs enabled (1 per chiplet) and the latency between cores on the same chiplet is lower than the latency to other chiplets. This allows AMD to target high end servers (up to 64 cores), low end (down to 16), workstations with threadripper (4 chiplets instead of 8), and high end desktops (2 chiplets instead of 8) with the same silicon. This allows them to spend less on fabs, R&D, etc because they can amortize the silicon over more products/volume. It also lets them bin them so chiplets with bad cores can still be sold. It's one of the things that lets AMD compete with the much larger volume Intel has, and do pretty well against numerous silicon designs Intel chips.
- throwaway4good 5y ago"Single-Chip Processors Have Reached Their Limits Announcements from XYZ and ABC prove that chiplets are the future, but interconnects remain a battleground" This could easily have been written 10 years ago, and I bet someone will write it in 10 years again. We need these really big chips with their big powerful cores because the nature the computing we do only changes very slowly towards being distributed and parallelizable and thus able to use a massive number of smaller but far more efficient cores.
- Dylan16807 5y agoYou're implying you can't put big powerful cores on chiplets but that's not true at all.
- lazide 5y agoHardly - performance/core hasn’t flatlined, but has not maintained the same growth over time (decades) in performance we’ve traditionally had. That’s the problem. So if you want better aggregate performance, more cores has been the plan for a decade+ now. FLOP/s per core or whatever other metric you choose to use. Previously it was possible to get 20-50% or more performance improvements even year to year for a core.
- Dylan16807 5y agoI wasn't talking about improvement at all. This was about big strong cores versus efficient cores, which is a tradeoff that always exists. You could choose between 20 strong cores or 48 efficient cores on the same die space across four chiplets, for example.
- alain94040 5y agoCorrect. Also known as Rent's rule. According to Wikipedia, it first was mentioned in the 1960s: https://en.wikipedia.org/wiki/Rent%27s_rule https://en.wikipedia.org/wiki/Rent%27s_rule
- marcodiego 5y agoMakes me remember the processor in the film terminator 2: https://gndn.files.wordpress.com/2016/04/shot00332.jpg https://gndn.files.wordpress.com/2016/04/shot00332.jpg
- WalterBright 5y agoI remember back in the 80's the limit was considered to be 64K RAM chips, because otherwise the defect rate would kill the yield. Of course, there's always the "make a 4 core chip. If one core doesn't work, sell it as a 3 core chip. And so on."
- fulafel 5y agoSome older stuff for reference: IBM POWER5 and POWER5+ (2004&2005) are MCM designs, had 2-4 CPU chips plus cache chips in same package. Link: https://en.wikipedia.org/wiki/POWER5 https://en.wikipedia.org/wiki/POWER5
- sliken 5y agoPentium pro from 1995 had two pieces of silicon in the package: https://en.wikipedia.org/wiki/Pentium_Pro https://en.wikipedia.org/wiki/Pentium_Pro
- h2odragon 5y agoPPros are quite hard to find now because the "gold scavengers" loved them. As i recall, at the peak in 2008, they were $100ea and more for the ceramic packages. All that interconnect was tiny gold wires, apparently.
- hinkley 5y agoI hope we are going to get back to a more asymmetric multi-processing arrangement in the near term where we abandon the fiction of a processor or two running the whole show with peripheral systems that have as little smarts as possible and promote them to at least second class citizens. These systems are much more powerful than when these abstractions were laid down, and at this point it feels like the difference between redundant storage on the box versus three feet away is more academic than anything else.
- wmf 5y agoThat kind of exists since most I/O devices have CPU cores in them, although usually hidden behind register-based interfaces. Apple has taken it a little further by using the same core everywhere and creating a standard IPC mechanism.
- gotaquestion 5y agoThe problem is AMP is very hard to program and debug. In embedded, one core is a scheduler and another is doing some real-time task (like arm BIG.little). In larger automotive heterogeneous compute platform, typically they are all treated as accelerators, or with bespoke Tier-1 integration (or like NVIDIA Xavier). And on top of that, OEMs always want to "reclaim" those spare cycles when the other AMP cores are underutilized, which is nigh impossible to do, so they fall back to symmetric MP. I think embedded is the only place for this to work right now. EDIT: I'm not an expert in this field but I have been asked to do work in this domain, and this narrow sampling is what I encountered, but I'd like to learn more about tooling and strategies for more generic AMP deployments.
- astrange 5y agoM1 is an AMP design, as is every iPhone SoC. It works well although you’ll be surprised if you try to run an SMP workload on every single core.
- gotaquestion 5y agoYes, I actually said this in my original post: Arm BIG.little works, which is what M1 is, and embedded works well, which are SoCs. It is much more difficult in AMP systems that aren't necessary only the same die, like automotive Tier 1 products.
- anonymousDan 5y agoSo is UCI-e a competitor/potential successor for something like Intel's QPI (or whatever they are using now)?
- narag 5y agoI hope somebody with relevant knowledge can answer this question, please: what % of the costs is "physical cost per unit" and what % is maintaining the I+D, factories, channels...? In other words, if a chip with 100x size (100x gates, etc.) made sense, would it cost 100x to produce or just 10x or just 2x? Edit: providing there wouldn't be additional design costs, just stacking current tech.
- mlyle 5y agoThere's many limiting factors... one is the reticle limit. But most fundamental is the defect density on wafers. If you have, say, 10 defects per wafer, and you have 1000 chips on it: odds are you get 990 good chips. If you have 10 chips on the wafer, you get 2-3 good chips per wafer. Of course, there's yield maximization strategies, like being able to turn off portions of the die if it's defective (for certain kinds of defects). For the upper limit, look at what Cerebras is doing with wafer scale. Then you get into related, crazy problems, like getting thousands of amperes into the circuit and cooling it.
- nightfly 5y agoI'm not an expert, or even an amateur, here but I defects are inevitable. So if you _need_ 100x the size without defects and one defect ruins the chip the cost might be 10000x to produce
- tonyarkles 5y agoIt's been a while since I've been out of that industry, but back around the 45nm days, one of the biggest concerns was yield. If you've got 100x the surface area, the probability of there being a manufacturing defect that wrecks the chip goes up. Now, you could probably get away with selectively disabling defective cores, but the chiplet idea seems, to me, like it would give you a lot more flexibility. As an example, let's say a chiplet i9 requires 8x flawless chips, and a chiplet Celeron requires 4 chips, but they're allowed to have defects in the cache because the Celeron is sold with a smaller cache anyway. In the "huge chip" case, you need the whole 8x area to be flawless, otherwise the chip gets binned as a Celeron. If the chiplet case, any single chip with a flaw can go into the Celeron bin, and 8 flawless ones can be assembled into a flawless CPU, and any defect ones go into the re-use bin. And if you end up with a flawed chip that can't be used at the smallest bin size, you're only tossing 1/4 or 1/8 of a CPU in the trash.
- bob1029 5y agoDespite the limitations apparently present in single chip/CPU systems, they can still provide an insane amount of performance if used properly. There are also many problems that are literally impossible to make faster or more correct than by simply running them on a single thread/processor/core/etc. There always will be forever and ever. This is not a "we lack the innovation" problem. It's an information-theoretic / causality problem you can demonstrate with actual math & physics. Does a future event's processing circumstances maybe depend on all events received up until now? If yes, congratulations. You now have a total ordering problem just like pretty much everyone else. Yes, you can cheat and say "well these pieces here and here dont have a hard dependency on each other", but its incredibly hard to get this shit right if you decide to go down that path. The most fundamental demon present in any distributed system is latency. The difference between L1 and a network hop in the same datacenter can add up very quickly. Again, for many classes of problems, there is simply no handwaving this away. You either wait the requisite # of microseconds for the synchronous ack to come back, or you hope your business doesnt care if john doe gets duplicated a few times in the database on a totally random basis.
- AnthonyMouse 5y agoThe alternative is speculative execution. If you can guess what the result is going to be, you can proceed to the next calculation and you get there faster if it turns out you were right. If you have parallel processors, you can stop guessing and just proceed under both assumptions concurrently and throw out the result that was wrong when you find out which one it was. This is going to be less efficient, but if your only concern is "make latency go down," it can beat waiting for the result or guessing wrong.
- tsimionescu 5y agoNot necessarily. There are problems you can't speed up even if you are given a literal infinity of processors - the problems in EXP for example (well, EXP - NP). Even for NP problems, the number of processors you need for a meaningful speed up grows proportionally to the size of the problem (assuming P!=NP).
- ksec 5y agoIs spectrum.ieee.org becoming another mainstream ( so to speak ) journalism where everything is dumbed down to basically Newspeak. The article is poorly written, the content is shallow and the headline is click bait.
- adhesive_wombat 5y agoThe IET magazine is the same, it's basically New Scientist but more expensive and you get letters after your name too. My favourite was the old PPARC Frontiers, was glossy but not breathless. But even then, these things never go into much detail.
- truth_seeker 5y agoA chip with Semi-FPGA as well as Semi-ASIC strategy could work. FPGA dev tools chain needs to improve.
- dahfizz 5y agoMaybe we will start to optimize our software instead of expecting accelerating CPUs to eat our bloat.
- sargstuff 5y agoTheoretically, only a problem if certfied turing complete.
- Animats 5y agoIs this a solution to a yield problem? Making physically bigger dies is no problem. Wafers are much larger than the individual dies. If the dies are just being laid out flat, there's no density gain. Multi-chip modules are nothing new. They've been used mostly when either there was a yield problem, or you wanted two different fab technologies. The latter is seen in some imagers and radars.
- pdimitar 5y agoWhat we need much more from here on are deterministic processors. There are awesome optimizers out there, including programs with genetic algorithms that find the best way to do a micro task X or Y (part of a much bigger program). IMO we have a ton of slightly-higher-hanging fruit that we can pick in terms of optimization but the relentless march of the X86 / X64 architecture obstructed that innovation. Might be time to look inwards and start working more on squeezing the CPUs we have right now for maximum performance.
- sargstuff 5y agoNot if C20 spec has anything to do with it. Program as OS, assign case selections in switch statement to threads (or distributed to other 'single-chip' cpu's. boundless networked single ASIC chip possibilities!
- ladyattis 5y agoI wonder if this will finally restart any research into non-Von Neumann architecture even for commercial uses like workstations and servers.