5 ms·
100x defect tolerance: How we solved the yield problem
- ChuckMcM 2y agoI think this is an important step, but it skips over that 'fault tolerant routing architecture' means you're spending die space on routes vs transistors. This is exactly analogous to using bits in your storage for error correcting vs storing data. That said, I think they do a great job of exploiting this technique to create a "larger"[1] chip. And like storage it benefits from every core is the same and you don't need to get to every core directly (pin limiting). In the early 2000's I was looking at a wafer scale startup that had the same idea but they were applying it to an FPGA architecture rather than a set of tensor units for LLMs. Nearly the exact same pitch, "we don't have to have all of our GLUs[2] work because the built in routing only uses the ones that are qualified." Xilinx was still aggressively suing people who put SERDES ports on FPGAs so they were pin limited overall but the idea is sound. While I continue to believe that many people are going to collectively lose trillions of dollars ultimately pursuing "AI" at this stage. I appreciate the the amount of money people are willing to put at risk here allow for folks to try these "out of the box" kinds of ideas. [1] It is physically more cores on a single die but the overall system is likely smaller, given the integration here. [2] "Generic Logic Unit" which was kind of an extended LUT with some block RAM and register support.
- __Joker 2y ago"While I continue to believe that many people are going to collectively lose trillions of dollars ultimately pursuing "AI" at this stage" Can you please explain more why you think so ? Thank you.
- mschuster91 2y agoIt's a hype cycle with many of the hypers and deciders having zero idea about what AI actually is and how it works. ChatGPT, while amazing, is at its core a token predictor, it cannot ever get to an AGI level that you'd assume to be competitive to a human, even most animals. And just as every other hype cycle, this one will crash down hard. The crypto crashes were bad enough but at least gamers got some very cheap GPUs out of all the failed crypto farms back then, but this time so much more money, particularly institutional money, is flowing around AI that we're looking at a repeat of Lehman's once people wake up and realize they've been scammed.
- KronisLV 2y ago> And just as every other hype cycle, this one will crash down hard. Isn't that an inherent problem with pretty much everything nowadays: crypto, blockchain, AI, even the likes of serverless and Kubernetes, or cloud and microservices in general. There's always some hype cycle where the people who are early benefit and a lot of people chasing the hype later lose when the reality of the actual limitations and the real non-inflated utility of each technology hits. And then, a while later, it all settles down. I don't think the current "AI" is special in any way, it's just that everyone tries to get rich (or benefit in other ways, as in the microservices example, where you still very much had a hype cycle) quick without caring about the actual details.
- anon373839 2y ago> I don't think the current "AI" is special in any way As someone who loves to pour ice water on AI hype, I have to say: you can't be serious. The current AI tech has opened up paths to develop applications that were impossible just a few years ago. Even if the tech freezes in place, I think it will yield substantial economic value in the coming years. It's very different from crypto, the main use case for which appears to be money laundering.
- carlmr 2y ago>It's very different from crypto, the main use case for which appears to be money laundering. Which has substantial economic value (for certain groups of people).
- lazide 2y agoAccording to this random estimate, black market economy alone in just the US is worth ~ $2 trillion/yr. [https://www.investopedia.com/terms/u/underground-economy.asp https://www.investopedia.com/terms/u/underground-economy.asp] Roughly 11-12% of GDP. In many countries, black+grey market is larger than the ‘white’ market. The US is notoriously ‘clean’ compared to most (probably top 10). Even in the US, if you suddenly stopped 10-12% of GDP we’re talking ‘great depression’ levels of economic pain. Honestly, the only reason Crypto isn’t bigger IMO is because there is such a large and established set of folks doing laundering in the ‘normal’ system, and those work well enough there is not nearly as much demand as you’d expect.
- ChuckMcM 2y agoI would guess you're not asking a serious question here but if you were feel free to contact me, it's why I put my email address in my profile.
- bigdict 2y agoWhy are you assuming bad faith?
- ChuckMcM 2y agoWhat gave you the impression I was assuming bad faith? It's off topic to the discussion (which is fine) but can be annoying in the middle of an HN thread.
- bigdict 2y ago> What gave you the impression I was assuming bad faith? You said "I would guess you're not asking a serious question here"
- ripped_britches 2y agoIt was a direct quote from your original comment
- bruce343434 2y agoYou brought it up...
- ossopite 2y agoWithout offering any opinion on its merits, if you think justifying this controversial claim is off topic, then so is the claim and you shouldn't have written it.
- kragen 2y agoYou said, "I would guess you're not asking a serious question here," which is to say, you were guessing that the question was asked in bad faith. Or, at any rate, you would, if for some reason the question came up, for example in deciding how to answer it. Which is what you were doing. That is to say, you did guess that it was asked in bad faith. Given the minimal amount of evidence available (12 words and a nickname "__Joker") I think it's reasonable to describe that guess as an assumption. Ergo, you were assuming bad faith.
- enragedcacti 2y agoAny thoughts on why they are disabling so many cores in their current product? I did some quick noodling based on the 46/970000 number and the only way I ended up close to 900,000 was by assuming that an entire row or column would be disabled if any core within it was faulty. But doing that gave me a ~6% yield as most trials had active core counts in the high 800,000s
- projektfu 2y agoThey did mention that they stash extra cores to enable the re-routing. Those extra cores are presumably unused when not routed in.
- enragedcacti 2y agoThat was my first thought but based on the rerouting graphic it seems like the extra cores would be one or two rows and columns around the border which would only account for ~4000 cores.
- projektfu 2y agoIf the system were broken down into more subdivisions internally, there would be more cores dedicated to replacement. It seems like it could be more difficult to reroute an entire row or column of cores on a wafer than a small block. Perhaps, also, they are building in heavy redundancy for POC and in the future will optimize the number of cores they expect to lose.
- ChuckMcM 2y agoI could guess that it helps with heat dissipation/management. But I don't know. That guess is from looking at the list of patents[1] they have. [1] https://patents.justia.com/assignee/cerebras-systems-inc https://patents.justia.com/assignee/cerebras-systems-inc
- girvo 2y ago> Xilinx was still aggressively suing people who put SERDES ports on FPGAs This so isn't important to your overall point, but where would I begin to look into this? Sounds fascinating!
- nroize 2y agoNot OP but I was curious too. Here's all I could find that seemed related: https://www.businesswire.com/news/home/20200121005582/en/Xilinx-Files-Patent-Infringement-Lawsuit-Against-Analog-Devices https://www.businesswire.com/news/home/20200121005582/en/Xil...
- ChuckMcM 2y agoWell this was the patent they were threatening with as I recall (https://patents.google.com/patent/US20030023912A1/en https://patents.google.com/patent/US20030023912A1/en) and there was this one too: https://patents.google.com/patent/US5576554A/en https://patents.google.com/patent/US5576554A/en Basically the "secret sauce" of the startup recruiting me was that they were going to do wafer scale FPGAs that could be tiled together to build arbitrarily complex systems like military phased array radars and such. All very hush hush but apparently they had recruited some key talent from Xilinx which was annoying Xilinx.
- deleted 2y ago[deleted]
- dogcomplex 2y agoOf course many people are going to collectively lose trillions, AI's a very highly hyped industry with people racing into it without an intellectual edge and any temporary achievement by any one company will be quickly replicated and undercut by another using the same tools. Economic success of the individuals swarming on a new technology is not a guarantee whatsoever, nor is it an indicator of the impact of the technology. Just like the dotcom bubble, AI is gonna hit, make a few companies stinking rich, and make the vast majority (of both AI-chasing and legacy) companies bankrupt. And it's gonna rewire the way everything else operates too.
- Melomomololo 2y ago[dead]
- idiotsecant 2y ago>it's gonna rewire the way everything else operates too. This is the part that I think a lot of very tech literate people don't seem to get. I see people all the time essentially saying 'AI is just autocomplete' or pointing out that some vaporware ai company is a scam so surely everyone is. A lot of it is scams and flash in the pan. But a few of them are going to transform our lives in ways we probably don't even anticipate yet, for good and bad.
- Retric 2y agoI’m not so sure it’s going to even do that much. People are currently happy to use LLM’s, but the outputs aren’t accurate and don’t seem to be improving quickly. A YouTuber watch regularly includes questions they asked Chat GPT and very single time there’s a detailed response in the comments showing how the output is wildly wrong from multiple mistakes. I suspect the backlash from disgruntled users is going to hit the industry hard and these models are still extremely expensive to keep updated.
- Thews 2y agoUsing function calls for correct answer lookup already practically eliminates this, it's not wide spread yet, but the ease of doing it is already practical for many. New models aren't being trained specifically on single answers which will only help. The expense for the larger models is something to be concerned about. Small models with function calls is already great, especially if you narrow down what they are being used for. Not seeing their utility is just a lack of imagination.
- wizzard0 2y agothis is an important reminder that all digital electronics is really analog but with good correction circuitry. and run-time cpu and memory error rates are always nonzero too, though orders of magnitude lower than chip yield rates
- nine_k 2y agoCPUs may be very digital inside, but DRAM and flash memory are highly analog, especially MLC flash. DDR4 even has a dedicated training mode [1], during which DRAM and the memory controller learn the quirks of particular data lines and adjust to them, in order to communicate reliably. [1]: https://www.systemverilog.io/design/ddr4-initialization-and-calibration https://www.systemverilog.io/design/ddr4-initialization-and-...
- ajb 2y agoSo they massively reduce the area lost to defects per wafer, from 361 to 2.2 square mm. But from the figures in this blog, this is massively outweighed by the fact that they only get 46222 sq mm useable area out of the wafer, as opposed to 56247 that the H100 gets - because they are using a single square die instead of filling the circular wafer with smaller square dies, they lose 10,025 sq mm! Not sure how that's a win. Unless the rest of the wafer is useable for some other customer?
- olejorgenb 2y agoIs the wafer itself so expensive? I assume they don't pattern the unused area, so the process should be quicker?
- yannyu 2y ago> I assume they don't pattern the unused area, so the process should be quicker? The primary driver of time and cost in the fabrication process is the number of layers for the wafers, not the surface area, since all wafers going through a given process are the same size. So you generally want to maximize the number of devices per wafer, because a large part of your costs will be calculated at the per-wafer level, not a per-device level.
- olejorgenb 2y agoYes, but my understanding is that the wafer is exposed in multiple steps, so there would still be less exposure steps? Probably insignificant compared to all the rest though. (Etching, moving the wafer, etc.) EDIT: to clarify - I mean the exposure of one single pattern/layer is done in multiple steps. (https://en.wikipedia.org/wiki/Photolithography#Projection https://en.wikipedia.org/wiki/Photolithography#Projection)
- yannyu 2y agoThe number of exposure steps would be unrelated to the (surface area) size of die/device that you're making. In fact, in semiconductor manufacturing you're typically trying to maximize the number of devices per wafer because it costs the same to manufacture 1 device with 10 layers vs 100 devices with 10 layers on the same wafer. This goes so far as to have companies or business units share wafers for prototyping runs so as to minimize cost per device (by maximizing output per wafer). Also, etching, moving, etc is all done on the entire wafer at the same time generally, via masks and baths. It's less of a pencil/stylus process, and more of a t-shirt silk-screening process.
- bee_rider 2y ago> Second, a cluster of defects could overwhelm fault tolerant areas and disable the whole chip. That’s an interesting point. In architecture class (which was basic and abstract so I’m sure Cerebras is doing something much more clever), we learned that defects cluster, but this is a good thing. A bunch of defects clustering on one core takes out the core, a bunch of defects not clustering could take out… a bunch of cores, maybe rendering the whole chip useless. I wonder why they don’t like clustering. I could imagine in a network of little cores, maybe enough defects clustered on the network could… sort of overwhelm it, maybe? Also I wonder how much they benefit from being on one giant wafer. It is definitely cool as hell. But could chiplets eat away at their advantage?
- IshKebab 2y agoTSMC also have a manufacturing process used by Tesla's Dojo where you can cut up the chips, throw away the defective ones, and then reassemble working ones into a sort of wafer scale device (5x5 chips for Dojo). Seems like a more logical design to me.
- mhh__ 2y agoAmazing. I clicked a button in the azure deployment menu today...
- ryao 2y agoI had been under the impression that Nvidia had done something similar here, but they did not talk about deploying the space saving design and instead only talked about the server rack where all of the chips on the mega wafer normally are. https://www.sportskeeda.com/gaming-tech/what-nvlink72-nvidia-ceo-jensen-huang-debuts-rtx-blackwell-gpu-shield-ces-2025 https://www.sportskeeda.com/gaming-tech/what-nvlink72-nvidia...
- bee_rider 2y agoIs this similar to a chiplet design? Chiplets have been a thing for a while, so I assume Cerebras avoided them on purpose.
- IshKebab 2y agoI don't think so - chiplets are much smaller and I think the process is different.
- iataiatax10 2y agoThe yield problem is not surprising they found a solution. Maybe they could elaborate more on the power distribution and dissipation problem?
- highfrequency 2y agoTo summarize: localize defect contamination to a very small unit size, by making the cores tiny and redundant. Analogous to a conglomerate wrapping each business vertical in a limited liability veil so that lawsuits and bankruptcy do not bring down the whole company. The smaller the subsidiaries, the less defect contamination but also the less scope for frictionless resource and information sharing.
- exabrial 2y agoI have a dumb question. Why isn't silicon sold in cubes instead of cylinders?
- bigmattystyles 2y agono matter how you orient a circle on a plane, it's the same
- amelius 2y agoThe silicon ingots have a rotating production process that results in cylinders, not bricks.
- exabrial 2y agofascinating, I figured it was something like that. maybe we should produce hexagonal, instead of square, chip designs
- kryptiskt 2y agoCrystalline silicon is produced with the Czochralski process (https://en.wikipedia.org/wiki/Czochralski_method https://en.wikipedia.org/wiki/Czochralski_method), which produces a round ingot. So you'd have to cut away perfectly fine silicon to make something squarish.
- NickHoff 2y agoNeat. What about power density? An H100 has a TDP of 700 watts (for the SXM5 version). With a die size of 814 mm^2 that's 0.86 W/mm^2. If the cerebras chip has the same power density, that means a cerebras TDP of 37.8 kW. That's a lot. Let's say you cover the whole die area of the chip with water 1 cm deep. How long would it take to boil the water starting from room temperature (20 degrees C)? amount of water = (die area of 46225 mm^2) * (1 cm deep) * (density of water) = 462 grams energy needed = (specific heat of water) * (80 kelvin difference) * (462 grams) = 154 kJ time = 154 kJ / 39.8 kW = 3.9 seconds This thing will boil (!) a centimeter of water in 4 seconds. A typical consumer water cooler radiator would reduce the temperature of the coolant water by only 10-15 C relative to ambient, and wouldn't like it (I presume) if you pass in boiling water. To use water cooling you'd need some extreme flow rate and a big rack of radiators, right? I don't really know. I'm not even sure if that would work. How do you cool a chip at this power density?
- lostlogin 2y agoIf rack mounted, you are ending up with something like a reverse power station. So why not use it as an energy source? Spin a turbine.
- sebzim4500 2y agoIf my very stale physics is accurate then even with perfect thermodynamic efficiency you would only recover about a third of the energy that you put into the chips.
- dylan604 2y ago1/3 > 0, so even if you don't get a $0 energy bill I'd venture that any company that could get 1/3 of energy bill would be happy
- bentcorner 2y agoI'm aware of the efficiency losses but I think it would be amusing to use that turbine to help power the machine generating the heat.
- bigmattystyles 2y agoWhen I was a kid, I used to get intel keychains with a die in acrylic - good job to whoever thought of that to sell the fully defective chips.
- dylan604 2y agowow, fancy with the acrylic. lots of places just place a chip (I'm more familiar with RAM sticks) on a keychain and call it a day.
- bigmattystyles 2y agothey're all over eBay, I just checked - the one I was thinking of, that I think I had is going for $150 - the things you get rid of....
- bradyd 2y agoElectronic Goldmine sells entire scrapped 200mm wafers for $15 or less https://theelectronicgoldmine.com/search?options%5Bprefix%5D=last&sort_by=price-descending&q=wafer https://theelectronicgoldmine.com/search?options%5Bprefix%5D...
- kragen 2y agoThose aren't just a chip; they're an epoxy package with a leadframe and a chip inside it. To put just a chip on a keychain, you'd have to drill a hole through it, which is difficult because silicon is so brittle—almost like drilling a hole in glass. Then, when someone put it onto a keyring, the keyring would form a lever that applies a massive force to the edge of the brittle hole, shattering the brittle silicon. Potting the chip in acrylic resin is a much cheaper solution that works better.
- Neywiny 2y agoUnderstanding that there's inherent bias by them being competitors of the other companies, but still this article seems to make some stretches. If you told me you had an 8% core defect rate reduced 100x, I'd assume you got to close to 99% enablement. The table at the end shows... Otherwise. They also keep flipping between cores, SMs, dies, and maybe other block sizes. At the end of the day I'm not very impressed. They seemingly have marginally better yields despite all that effort.
- sfink 2y agoI think you're missing the point. The comparison is not between 93% and 92%. The comparison is between what they're getting (93%) and what you'd get if you scaled up the usual process to the core size they're using (0%). They are doing something different (namely: a ~whole wafer chip) that isn't possible without massively boosting the intra-chip redundancy. (The usual process stops working once you no longer have any extra dies to discard.) > Despite having built the world’s largest chip, we enable 93% of our silicon area, which is higher than the leading GPU today. The important part is building the largest chip. The icing on the top is that the enablement is not lower. Which it would be without the routing-to-spare-cores magic sauce. And the differing terminology is because they're talking about differing things? You could call an SM a core, but it kind of contains (heterogeneous) cores itself. (I've no idea whether intra-SM cores can be redundant to boost yield.) A die is the part you break off and build a computer out of, it may contain a bunch of cores, a wafer can be broken up into multiple dies but for Cerebras it isn't. If NVIDIA were to go and build a whole-wafer die, they'd do something similar. But Cerebras did it and got it to work. NVIDIA hasn't gotten into that space yet, so there's no point in building a product that you can't sell to a consumer or even a data center that isn't built around that exact product (or to contain a Balrog).
- fspeech 2y agoThere is nothing inherently good about wafer scale. It's actually harder to dissipate heat and enable hybrid bonding with DRAM. So the gp is entirely correct that you need to actually show higher silicon utilization to be even considered as being something worthwhile.
- wendyshu 2y agoWhat's yield?
- anonymousDan 2y agoVery interesting. Am I correct in saying that fault tolerance here is with respect to 'static' errors that occur during manufacturing and are straightforward to detect before reaching the customer? Or can these failures potentially occur later on (and be tolerated) during the normal life of the chip?
- abrookewood 2y agoLooking at the H100 on the left, why is the chip yield (72) based on a circular layout/constraint? Why do they discard all of the other chips that fall outside the circle?
- flumpcakes 2y agoBecause the circle is the physical silicon. Any chips that fall outside the circle are only part of a full chip. They will be physically missing half the chip.
- donavanm 2y agoAFAIK all wafer ingots are cylinders, which means the wafers themselves are a circular cross section. So manufacturing is binpacking rectangles in to a circle. Plus different effects/defects in the chips based on the distance from the edge of the wafer. So I believe its the opposite: why are they representing the larger square and implying lower yield off the wafer in space that doesnt practically exist?
- therealcamino 2y agoThat's just the shape of the wafer. I don't know why the diagram continued the grid outside it.
- ryao 2y ago> Take the Nvidia H100 – a massive GPU weighing in at 814mm2. Traditionally this chip would be very difficult to yield economically. But since its cores (SMs) are fault tolerant, a manufacturing defect does not knock out the entire product. The chip physically has 144 SMs but the commercialized product only has 132 SMs active. This means the chip could suffer numerous defects across 12 SMs and still be sold as a flagship part. Fault tolerance seems to be the wrong term to use here. If I wrote this, I would have written redundant.
- jjk166 2y agoRedundant cores lead to a fault tolerant chip.
- ryao 2y agoECC memory is fault tolerant. It repairs issues on the fly without disabling hardware. This on the other hand is merely redundant to handle manufacturing defects. If they make a mistake and ship a bad core that malfunctions at runtime, it is not going to tolerate that.
- jjk166 2y agoRedundancy is a method of providing fault tolerance, the existence of other methods doesn't make it less fault tolerant. Nothing is tolerant to all possible faults. Fault tolerance refers to being able to tolerate specific types of faults under specific conditions. Fault tolerant is the proper term for this.
- ryao 2y agoI think it would have been better to write redundant. It is more specific.
- gunalx 2y agoMy biggest question is who are the buyers?
- asdasdsddd 2y agomostly 1 ai company in the middle east last I heard
- bcatanzaro 2y agoThis is a strange blog post. Their tables say: Cerebras yields 46225 * .93 = 43000 square millimeters per wafer NVIDIA yields 58608 * .92 = 54000 square millimeters per wafer I don't know if their numbers are correct but it is a strange thing for a startup to brag that it is worse than a big company at something important.
- saulpw 2y agoBeing within striking distance of SOTA while using orders of magnitude fewer resources is worth bragging about.
- RecycledEle 2y agoIIRC, it was Carl Bruggeman's IPSA Thesis that showed us how to laser out bad cores.
- oksurewhynot 2y agoI live in a small city/large town that has a large number of craft breweries. I always marveled at how these small operations were able to churn out so many different varieties. Turns out they are actually trying to make their few core recipes but the yield is so low they market the less consistent results as...all that variety I was so impressed with.
- tweetpeekai 2y ago[dead]
- trhway 2y ago56K mm2 vs 46K mm2. I wonder why they wouldn’t use the smart routing/etc to use more fitting shape than square and thus use more of the wafer.
- ilaksh 2y agoI assume people are aware, but Cerebras has a web demo and API which is open to try and it is 2000 tokens per second for Llama 3.3 70b and 1000 tokens per second for Llama 3.1 405b. https://cerebras.ai/inference https://cerebras.ai/inference
- Fokamul 2y agoAnyone has some picture how it is looks like inside these servers?
- hoseja 2y agoWhy square chip? Make it an octagon or something.
- aurareturn 2y agoBear case on Cerebras: https://irrationalanalysis.substack.com/p/cerebras-cbrso-equity-research-report https://irrationalanalysis.substack.com/p/cerebras-cbrso-equ... Note: This author is heavily invested in Nvidia.
- larsrc 2y agoHow do these much smaller cores compare in computing power to the bigger ones? They seem to implicitly claim that a core is a core is a core, but surely one gets something extra out of the much bigger one?
- jstrong 2y agoI would like a workstation with 900k cores. lmk when these things are on ebay.
- aaroninsf 2y agoThe number of people ITT this thread who have absorbed the world-weary AI-is-a-bubble skepticism... I'm just gonna say, with serene certainty, the economic order we inhabit going through phase change is certain. From certain myopic perspectives we can shoehorn that into a narrative of cyclical patterns in the tech industry or financial markets etc etc. This is not going to be that. No more than the transformation of American retail can be shoehorned to kind of look like it used if you don't know anything at all about what contemporary international trade and logistics and oligopoly actually mean in terms of what is coming into your home from where and why it is or isn't cheap. Where we'll be in 10, 20, years is literally unimaginable today; and trying to navigate that wrt traditional landmarks... oof.
- lofaszvanitt 2y agoA well written, easy to understand article.
- TowerTall 2y agoEver heard the old joke story about an American buyer told the Japanese manufacture how many incorrectly made bolts were acceptable per lot of a thousand bolts? Maybe 2 or 3 in 1,000? So the Japanese didn't have any incorrectly made bolts in their manufacturing process so they just added two or three bad ones to every batch to please the Americans.
- ashvardanian 2y agoThe AMD comparison may not be accurate. The 96 core AMD CPU takes multiple such dies (eight if I remember correctly) and separate IO chiplets. The total surface area listed should be much larger.