7 ms·
Chip aging becomes design problem (2018)
- Lind5 5y agoalso relevant & more recent https://semiengineering.com/reliability-concerns-shift-left-into-chip-design/ https://semiengineering.com/reliability-concerns-shift-left-...
- trasz 5y agoSo, how serious it is for current and upcoming hardware, realistically?
- monocasa 5y agoA real problem for smaller nodes, 7nm and under starts really cutting into expected lifetime.
- Animats 5y agoThis is a big problem for automotive. The average age of cars in the US is over 12 years now. Which is a technical achievement. It was around 6 in the 1970s. Longer for trucks. Automotive electronics design lives used to be much longer. I happen to know that the design life for the Ford EEC IV engine control unit from the 1980s was 30 years. That's been exceeded in the field; many 1980s Ford trucks with that unit are still running. There are now cars running electronics that's way overkill in performance, probably at the cost of lifetime. Unreal Engine in the dashboard is a thing. Realistically, you need 20 years of life in automotive electronics.
- versteegen 5y agoIt would be fantastic, market changing, to know what cars or other consumer goods were actually designed for, but I don't have hope for that. Honestly, I wouldn't buy a car if I knew its electronics were only designed to last 20 years. I wouldn't expect it to be reliable as it approaches that age. Mind you, I drive a 1994 Toyota Corolla, which has needed about $50 of repairs in the last decade, so I have high standards.
- Robotbeat 5y ago2-3 years for a consumer device and 10 for a telecommunications device s em both massively too short by a factor of 2 or more. Do we really want our devices to stop working physically, at the chip level, in less than a decade? Some people still play on game consoles decades old; these are most certainly consumer devices. I’d much rather we over-design stuff to last decades, at least at the chip level where overedesigning is super cheap. Solid state electronics was supposed to mean longer life for everything. No tubes to burn out, no mechanical parts to wear out. It would suck so bad if we nickel and dimed what should be fundamentally physically robust devices to last for much less time than complicated mechanical devices from the past.
- nashashmi 5y agoThey don't stop working after 2_3 years. They just don't work as fast. The electronics industry wants to promote change of products. so they do this. Or they don't work on long term durability. My take: for as long as recycling performance is terrible, this should be a no no. Social movements should start demanding better products.
- nickff 5y ago>"Do we really want our devices to stop working physically, at the chip level, in less than a decade? Some people still play on game consoles decades old; these are most certainly consumer devices." Many components in your consumer devices age; my greatest concern is electrolytic capacitors (specifically the tantalum ones). I think many game consoles and computers with switched-mode power supplies are unlikely to last decades. If you want a simple way to get your devices to last longer, I suggest that you pick ones with external (brick) power supplies, as the SMPS is likely to be one major cause of failures. >"I’d much rather we over-design stuff to last decades, at least at the chip level where overedesigning is super cheap." I am not sure that people are willing to pay significantly more for these longer-lasting devices, and all that extra effort will be wasted if the devices are scrapped prematurely. It may be more (environmentally) efficient to simply replace the devices upon failure rather than over-designing them.
- ece 5y ago
- ohazi 5y agoHow long should I expect to be able to use a new 5nm CPU (at reasonable temperatures) before these issues are likely to make it fail? All of the desktop/laptop CPUs that I currently use are 14nm, and I think the oldest is around 7 years old and still working fine. In the past I've tended to use personal machines for around a decade, and I don't really have any desire to move to a shorter cycle. Better battery life is great, but most things are already plenty thin and fast.
- piyh 5y agoIt's all about duty cycle. Idle browsing is not taxing.
- tester756 5y agoReminds me of [Cores that don't count]https://research.google/pubs/pub50337/ https://research.google/pubs/pub50337/ >We are accustomed to thinking of computers as fail-stop, especially the cores that execute instructions, and most system software implicitly relies on that assumption. During most of the VLSI era, processors that passed manufacturing tests and were operated within specifications have insulated us from this fiction. As fabrication pushes towards smaller feature sizes and more elaborate computational structures, and as increasingly specialized instruction-silicon pairings are introduced to improve performance, we have observed ephemeral computational errors that were not detected during manufacturing tests. These defects cannot always be mitigated by techniques such as microcode updates, and may be correlated to specific components within the processor, allowing small code changes to effect large shifts in reliability. Worse, these failures are often "silent'': the only symptom is an erroneous computation. >We refer to a core that develops such behavior as "mercurial.'' Mercurial cores are extremely rare, but in a large fleet of servers we can observe the correlated disruption they cause, often enough to see them as a distinct problem -- one that will require collaboration between hardware designers, processor vendors, and systems software architects. >We have observed various kinds of symptoms caused by mercurial cores. > Violations of lock semantics leading to application data corruption and crashes. > Data corruptions exhibited by various load, store, vector, > and coherence operations. > A deterministic AES mis-computation, which was “self inverting”: encrypting and decrypting on the same core yielded the identity function, but decryption elsewhere yielded gibberish. > Corruption affecting garbage collection, in a storage system, causing live data to be lost. > Database index corruption leading to some queries, depending on which replica (core) serves them, being non deterministically corrupted. > Repeated bit-flips in strings, at a particular bit position (which stuck out as unlikely to be coding bugs). > Corruption of kernel state resulting in process and kernel crashes and application malfunctions ___________ >Not all mercurial-core screening can be done before CPUs are put into service – first, because some cores only become defective after considerable time has passed,
- ccbccccbbcccbb 5y agoPlanned obsolescence. This may well end up with CPU vendors adopting food storage cant. "Ignel i11-23017K, MFD: 2025.05.01, EXP: 2026.09.03, consume within 3 months of first power-on".
- justicezyx 5y agoSemi conductor newbie here. Reading this, I am wondering what failures in the transistor (field effect transistor) can be tolerated by the circuit? I forgot all the analog and digital circuitry classes in college. I don't recall there is any redundancy in the circuit design. Is redundancy built in the circuit design therefore the chip?
- ted_dunning 5y agoFor some circuits, there can be substantial redundancy. This includes memory, disks and networks. Sometimes that redundancy is in hardware, sometimes software and sometimes a mix. For other hardware, there is almost the of redundancy. In those parts of the system, you depend on multiple components all working correctly with no chance of detecting errors at the circuit level. This means that if you have 10 components with expected life of 6-20 years (with a mean of 10) you can expect an actual life of about 6, not the mean life at all. The weakest link and all that.
- GeekyBear 5y agoI wondered it this might be a reason for Intel disabling AVX-512 on it's most recent chips. >Electromigration is the movement of atoms based on the flow of current through a material. If the current density is high enough, the heat dissipated within the material will repeatedly break atoms from the structure and move them. This will create both ‘vacancies’ and ‘deposits’. The vacancies can grow and eventually break circuit connections resulting in open-circuits, while the deposits can grow and eventually close circuit connections resulting in short-circuit... In Black’s equation, which is used to compute the mean time to failure of metal lines, the temperature of the conductor appears in the exponent, i.e. it strongly affects the MTTF of the interconnect https://www.synopsys.com/glossary/what-is-electromigration.html https://www.synopsys.com/glossary/what-is-electromigration.h... Intel's new chips already run hot and running AVX-512 instructions has required increasing the voltage. >One of the big takeaways from our initial Core i7-11700K review was the power consumption under AVX-512 modes, as well as the high temperatures. Even with the latest microcode updates, both of our Core i9 parts draw lots of power. The Core i9-11900K in our test peaks up to 296 W, showing temperatures of 104ºC, before coming back down to ~230 W and dropping to 4.5 GHz. There are a number of ways to report CPU temperature. We can either take the instantaneous value of a singular spot of the silicon while it’s currently going through a high-current density event, like compute, or we can consider the CPU as a whole with all of its thermal sensors. While the overall CPU might accept operating temperatures of 105ºC, individual elements of the core might actually reach 125ºC instantaneously. So what is the correct value, and what is safe? https://www.anandtech.com/show/16495/intel-rocket-lake-14nm-review-11900k-11700k-11600k/5 https://www.anandtech.com/show/16495/intel-rocket-lake-14nm-...
- ggm 5y agoAnyone else holding a Sun Ultra-5 or Ultra-10 where the MAC address has zero'ed out? Sometimes, the reason you can't be bootstrapped is as simple as a soldered-on battery backup (there's a hex boot load sequence to give your host a self-assigned MAC and get over this problem)