4 ms·
Chips are not designed to have fault tolerance. The typical solution on large chips today is to have large number of identical units, and use fuses to deactivat
by guepe 6y ago
Chips are not designed to have fault tolerance. The typical solution on large chips today is to have large number of identical units, and use fuses to deactivate some when they are faulty. This is called binning and how all manufacturers build their line of products.
Design and validation of fault-tolerant chips is MUCH more involved, which slows down next-gen products a lot. It's in general better to simply use today's process ASAP, with no or little fault tolerance.
The exception is wafer-scale chips, which are chips that span multiple "reticles", i.e. the size of the masks used to manufacture normal chips.
I have a Ph.D. thesis linked to building a wafer scale chip (ala cerebras, although 10 years ago). And cerebras AFAIK is the FIRST commercial product using wafer-scale approaches. Maybe we will see more products ? I have doubts...
Note also that defect density varies depending on location of the chip on the wafer. There are areas with more defects on average, so it's not as simple as 1 number.
- baybal2 6y ago> And cerebras AFAIK is the FIRST commercial product using wafer-scale approaches. Amdahl?
- mlyle 6y agoI thought about mentioning Gene Amdahl/Trilogy in my response. But they never actually got to market.
- mlyle 6y ago> This is called binning Binning is any kind of sorting by silicon quality. It can refer to sorting by the number of "good cores" you get, but it has usually meant sorting by speed grade or other characteristics. > and how all manufacturers build their line of products. Having identical functional units you disable is a great yield maximization strategy if it fits your product. Most semiconductor products are not in this category-- it's mostly just multicore CPUs and GPUs that have a lot of identical cores and relatively small uncore.
- guepe 6y agoYes I agree with your comments, I over simplified. However, regarding "most semi", it depends if we are talking about market value vs # chips. The CPU/GPU market is really much larger in $$$ than low-power embedded chips. Although it's changing again with AI inference...
- brennanpeterson 6y agoAgreed on binning, but all DRAM and all NAND are fault tolerant. Between all memory, GPU and GPU-like, and CPU....I have all but mobile, or about 70% of the market.
- mlyle 6y agoYes, memories too, are a good mention. > I have all but mobile, or about 70% of the market. In dollars, maybe, but there a whole lot of units of things that you cannot build a meaningful deactivation strategy on.
- brennanpeterson 6y agoSure! It is a huge mix. ASICs don't, mobile cannot, and you can't really just ignore. But there are also other tricks, like redundant vias, or SRAM spares, which do make 'yield' less listed by defects than it might first appear.
- jhallenworld 6y agoFPGAs too, though I don't know how widespread this technique is: https://www.eetimes.com/altera-uses-redundancy-to-boost-yields-in-high-density-plds/# https://www.eetimes.com/altera-uses-redundancy-to-boost-yiel...
- rbanffy 6y ago> It can refer to sorting by the number of "good cores" you get, but it has usually meant sorting by speed grade or other characteristics. There are multiple "bins" between "good core" and "failed core". The first one runs flawlessly at the nominal speed while the second one can't run within the envelope defined. There are cores that can run overclocked, cores that will fail above a certain frequency, and cores that are so broken they won't work on any clock. That's why, for instance, some Cell processors left the factory with 6 SPUs and Sun's first Niagaras had SKUs with only a small number of cores.
- jules 6y agoOn what scale is the deactivation done? Products are differentiated by number of cores, but do they deactivate parts on a more fine grained scale?
- mlyle 6y agoSometimes lumps of cache or memory chunks. In the past, it was an entire FPU.