3 ms·
As far as I understand, chips are designed with many redundant pathways for exactly this reason. I think its a rather important part of the design process, but
by jflatow 11y ago
As far as I understand, chips are designed with many redundant pathways for exactly this reason. I think its a rather important part of the design process, but someone who knows better should probably say.
- valarauca1 11y agoNot exactly. Reduancy is designed into chip hardware not to account for cosmic rays striking the wrong trace, but for manufacturing defects. For example Nvidia runs it's GTX 980 and GTX 970 production at the same time. The only difference is GTX970's can have up to 2 of their compute units non-functional. This is very commonly done in the industry. If you remember Phenom Dual, Tri, and Quadcores. Which were the same chip, just it was expected that only 5% of produced chips would be fully featured quad cores, the rest would be sold as other core counts. This was done with 27xx series i5, which were 34xx core i7's but with hyper threading disabled due to issues with yields on dye shrinks. If a single transistor fails, normally the whole thing dies.
- bri3d 11y agoThey're designed for manufacturing defects - this process is called DFM (Design For Manufacturing) and usually revolves around performing a Critical Area Analysis for defects of a given size. Design software is used to determine what would happen if a speck of dust of a certain diameter landed on the die in given locations. Then, the critical areas are spaced out and moved around to attempt to balance design constraints with yield. For large, expensive parts or parts in which a single common defect could easily blow the whole yield (for example DRAM, especially when embedded, CPUs with lots of cache or cores, and so on), regions (or rows and columns of memory) are generally fused off so that if one specific region fails qualification, it can be disabled without discarding the whole chip. This is the source of most 3-core CPUs, as well as the difference between most models in a single CPU family (they're often binned off based on how much of their L2 cache actually works). However, once parts are manufactured and qualified, they're pretty much done. Some hardware has BISR (Built In Self Repair) but as far as I know it's not particularly common outside of DRAM.
- revelation 11y agoBuilt-in-self-repair is a key feature in the old spinning rust hard drives and particular the new flash memory.
- ajross 11y agoIn neither case is the device "repairing" a failure like the ones posited here, though. Storage devices "repair" failures by detecting them and recovering the data using some form of ECC, and then rewriting it to a location that has not failed. No one has yet figured out a way to have a shorted polysilicon feature un-short itself in situ. :)