9 ms·
CPU reliability – Linus Torvalds (2007)
- bcoates 13y agoThread context: https://lkml.org/lkml/2007/5/11/179 https://lkml.org/lkml/2007/5/11/179
- pedrocr 13y agoIt would be awesome if companies like Google would calculate MTBF statistics on components. They've done it for disks and it would be great to extend it to CPUs and memory modules. They're probably in a better position than even Intel to calculate these things with precision.
- zebra 13y agoI'm almost sure that the components without moving parts will become technologically obsolete long before they start to fail. When I buy used laptop I always change the HDD, the DVD and its reliability jumps sharply up.
- jfim 13y agoEven things with moving parts, it would be nice to know that model X of brand Y has a MTBF of 4.5 years, but hunting the same model X 4.5 years later isn't likely to yield the exact same hardware but some later revision of the same specced hardware.
- pedrocr 13y agoThat may very well be true on average but I'd bet there are plenty of CPUs and memory modules that fail in the first year of usage for example. After all CPUs are tested and sorted into high/low performance parts, so sample variation itself would be enough to generate some early failures. As a consumer it's hard enough to keep up with what's reliable in hard drives. Keeping the manufacturers honest with good stats for the most common parts would be great.
- superuser2 13y agoI've seen a lot of failed laptop motherboards.
- wging 13y agoThey might very well do this already. But it'd seem they're disincentivized to make this information public... do they publish the disk information?
- pedrocr 13y agoThey've published some stuff on disks: https://static.googleusercontent.com/external_content/untrusted_dlcp/research.google.com/en/us/archive/disk_failures.pdf https://static.googleusercontent.com/external_content/untrus... The backblaze guys also have a big data set: http://blog.backblaze.com/2013/11/12/how-long-do-disk-drives-last/ http://blog.backblaze.com/2013/11/12/how-long-do-disk-drives...
- amscanne 13y agoInteresting relevant paper: http://www.cs.cmu.edu/~bianca/fast07.pdf http://www.cs.cmu.edu/~bianca/fast07.pdf Table 3 suggests that there are data sets that include all components (CPU, memory, power supplies, etc.).
- avn2109 13y agoData of the sort shown in Table 3 certainly exists. However, it's often inaccessible to the public because companies tend to treat it as a trade secret. Any large company that makes things employs a bunch of reliability engineers, who are usually EE's or ME's who make Weibull plots and bathtub curves all day (to set the warranty duration, mostly). These guys have all the data you could ever want on this topic, but they're not sharing. Especially at Intel.
- rkangel 13y agoThey'd have to be careful with how they quoted the numbers though. As Linus accurately points out, MTBF varies wildly depending on the usage pattern. If you want to quote it in a unit of time, e.g. "years", then you have to specify the usage the part has been under, which will be very different for a server part compared to a desktop part. You could quote it per instruction or equivalent, I suppose, taking into account how hard the component is used, but even that isn't perfect.
- bobdvb 13y agoMTBF is quite well tied to a contract purchasing when you buy large quantities of components for manufacturing. Intel don't know where people are going to use their products and the range of environments they get exposed to can't be easily equated for. I know Intel has been working for some time on the idea of high temperature data centres, this will impact the MTBF of all components but you can always calculate the cost of the losses vs the cost of the cooling: http://www.datacenterdynamics.com/focus/archive/2012/08/intel-pushes-temperature-data-center http://www.datacenterdynamics.com/focus/archive/2012/08/inte...
- bcoates 13y agoHere's one for RAM in servers, from Google: http://news.cnet.com/8301-30685_3-10370026-264.html http://news.cnet.com/8301-30685_3-10370026-264.html Found it here, which also goes into some testing the Guild Wars guys did on their population of gamer PCs: http://www.codeofhonor.com/blog/whose-bug-is-this-anyway http://www.codeofhonor.com/blog/whose-bug-is-this-anyway (scroll down to "Your computer is broken", around 1% of the systems they tested failed a CPU-to-RAM consistency stress test) Both of them indicate intermittently defective components in running systems are way more common than anybody assumes.
- raverbashing 13y agoThis is very interesting But when they say "Memory Error", even though it's something detected/corrected by ECC I'm not sure we can say 'the memory is defective' It may be a combination of the conditions of power/load/data/time since last refresh and variance between modules. Since Google appears not to show all the data they have, we probably are not going to get that from them though :/
- leokun 13y agoNice thing about the cloud is that someone else is worrying about this for you.
- heaviside 13y agoThis study by Microsoft Research is interesting: "Cycles, Cells and Platters: An Empirical Analysis of Hardware Failures on a Million Consumer PCs" http://research.microsoft.com/apps/pubs/default.aspx?id=144888 http://research.microsoft.com/apps/pubs/default.aspx?id=1448...
- sytelus 13y agoIf MTBF is such a big issue then would it be ever possible to build space craft that travels across the stars and still has ability communicate? I guess hats off to designers of Voyager and other spacecrafts whose MTBF seems to have crossed 36+ years for many components including CPU and power supply. But for inter-steller crafts that MTBF seems VERY low. And, seriously, MTBF of 5 years seems to be joke for desktop when lot of mechanical components with moving parts actually lasts longer.
- seiji 13y agoYou can fake whole system reliability by incorporating redundant internal systems.
- hkmurakami 13y agoThat makes me wonder... Linus refers to this as well, but how much of the 36+ years can be attributed to the components actually being turned off? Also, I'd imagine that space craft components are of an entirely different category of components that the off the shelf computing variety.
- greenyoda 13y agoSpacecraft are not built out of the same grade of components as consumer and commercial hardware.
- eck 13y agoHe's talking about desktop/server CPUs where people care about performance. If you don't care so much about performance, you can increase the transistor sizes, reduce the clock speed, and achieve totally insane MTBF... as space-rated hardware tends to do. Kind of like how server CPUs are underclocked to increase MTBF, but more so.
- nwh 13y agoSpacecraft and rovers use ridiculously armoured, redundant systems to get past the fact that they would fail quite regularly in such a hostile environment. The Curiosity rover in 2001 uses what would normally be quite an outdated 132Mhz CPU that's been specially shielded to achieve the reliability the program needs; even then there's two redundant systems that do health checks on one other to avoid bit flips. Even with all of that, they're running on only one CPU and trying to diagnose why the first one failed. It's probably not fair to compare the MTBF of specialised hardware to the $35 CPU I bought at the retailer down the street either, the RAD750 processors in Curiosity cost almost a quarter of a million dollars each. http://en.wikipedia.org/wiki/Comparison_of_embedded_computer_systems_on_board_the_Mars_rovers http://en.wikipedia.org/wiki/Comparison_of_embedded_computer... http://en.wikipedia.org/wiki/Curiosity_rover#Specifications http://en.wikipedia.org/wiki/Curiosity_rover#Specifications http://en.wikipedia.org/wiki/Radiation_hardening#Radiation-hardening_techniques http://en.wikipedia.org/wiki/Radiation_hardening#Radiation-h... Though that said, Voyager is still happy running on it's 8064 words of 16 bit RAM, which is something.
- raverbashing 13y ago(Conventional) Solid state devices are very hard to fail - exception: flash memory Apart from electron migration issues and failures by excess (voltage/temperature), they're pretty long lasting Much easier to have a failure because of something else: capacitors failing, oxidation or mechanical failure (for example, because of thermal expansion/contraction) I've seen people complaining about a dead CPU but I can't find it right now
- soundsop 13y agoYou are correct. I want to clarify that the failure process is electromigration, not electron migration. It is caused by electrons but it is ions in metal that migrate. Wikipedia has a good description: https://en.wikipedia.org/wiki/Electromigration https://en.wikipedia.org/wiki/Electromigration. I design integrated circuits and one of the constraints in selecting the width of wires is to make sure that the maximum current density is below the electromigration threshold.
- zymhan 13y agoI actually just returned my CPU (Phenom II X4) to AMD, and they've replaced it, but they didn't say exactly why it died. I've asked them for more details, hopefully they can tell me. Overall though, given with how many computers I've worked with, CPU failures still seem rarer than Memory, Disk, Mobo, or Graphics failures. Of course it ends up being the CPU in _my_ computer that fails -.-
- raverbashing 13y agoInteresting. How long did it work for before it died? There is still some variance in silicon, so yours may have had a defect that manifested itself after some time, I'm not sure they evaluate returned defective chips to see what happened (and if this is public info) Also, the packaging is extremely complex and prone to the same kind of defects as other PCBs in the system.
- Zardoz84 13y agoI can say that the Z80 if my ZX Spectrum keep working since 1984... Or some old K6-2 300 was working this last year...
- caf 13y agoFor how much of the time since 1984 would you imagine that your Z80 has been on and running?
- AnonNo15 13y agoI'd like to through my experience: I was in charge of 300+ x86 rack servers and around 50 desktops for 3 years and never seen a single CPU fail, even old Pentium 4 with dusty fans. Disk failures are very common, followed by much rarer RAM chips and motherboards failures. I suspect server chips are rated for 10-15 years average lifespan
- dspeyer 13y agoIt doesn't seem worth it for Intel to measure MTBF. By the time they got good numbers for a specific chip, they'd be trying to sell its successor.
- gilgoomesh 13y agoLong term failure rates are not usually measured in realtime but in deliberately heat elevated environments which simulate many years of stresses in a few months. This work is essential to ensure design decisions they've made don't accidentally cause their chips to fail after 2 years (which might be outside warranty lifetime but would still result in class action law suits and horrible publicity).
- williadc 13y agoIntel guarantees their consumer CPUs for 3 years. http://www.intel.com/support/processors/sb/cs-020033.htm http://www.intel.com/support/processors/sb/cs-020033.htm
- zxcdw 13y agoI don't work in environment where I get to deal with hardware failures, so pardon my ignorance, but has anyone seen a failed CPU piece which has failed during normal operation? I am under an impression that it is very rare for a CPU itself to fail so that it would need to be replaced. The only times I've even heard about failing CPUs has been if they've been overclocked or insufficiently cooled(add in overvolting, and you get both :)) or physical damage during mounting/unmounting or otherwise handling hardware. And even then the failure has usually been elsewhere than the CPU itself. Of course I am not saying it'd be unheard of, but for me frankly, right now it is.
- zhemao 13y agoI think the MTBF is generally longer than people would normally go without replacing their CPU. Also, CPUs are generally designed to degrade more gracefully. For instance, they may have circuity that scales the frequency down as delays get longer. Also, in multicore CPUs, there are generally some spare cores that will get swapped in if a previously in-use core breaks.
- anon_cownerd 13y ago> Also, in multicore CPUs, there are generally some spare cores that will get swapped in if a previously in-use core breaks. That sounds like a huge cost to bear. Looking at e.g. a Haswell die photo [1], there are just four physical cores present for a four-core part. With that die area per core, you would take a ~15-20% area hit (that translates to 15-20% cost) just to have a spare core in case one failed some years later. I have heard of manufacturers selling otherwise "defective" parts where a core or cache slice has a defect by relabeling as a part with fewer cores. But that's a manufacture-time decision, not a dynamic reconfiguration in the field. [1] http://cdn2.wccftech.com/wp-content/uploads/2013/05/Intel-Haswell-Core.jpg http://cdn2.wccftech.com/wp-content/uploads/2013/05/Intel-Ha...
- ssafejava 13y agoI think the GP was on the right track, but somewhat confused. They obviously don't do this on all models, but some dual-core models are disabled quad-cores. Remember the Athlon X3? That was a binned chip that usually was created from X4s with a broken core. Most buyers didn't mind, and some of them got lucky and were able to re-enable the disabled core. It seems like the GP might be suggesting that a quad-core CPU will swap in another core when one dies. That doesn't happen. But the binning process allows them to still sell slightly defective silicon with disabled parts (cores, cache), which saves money. On a related note, a lot of GPUs actually do have a few dozen execution units that are disabled by default and can be swapped in after stress testing at the factory. I believe some can even do that in the wild, but I could be wrong.
- Taniwha 13y agoso not even mentioned here is metastability - basically signals that cross clock domains within traditional clocked logic where the clocks are not carefully organized to be multiples of each other can end up being sampled just as they change - the result is a value inside of a flip-flop that's neither a 1 or a 0 - sometimes an analog value somewhere in between, sometimes an oscillating mess at some unknown frequency - worst worst case this unknown bad value can end up propagating into a chip causing havoc, a buzzing mess of chaos. In the real world this doesn't happen very often and there are techniques to mitigate it when it does (usually at a performance or latency cost) - core CPUs are probably safe, they're all one clock but display controllers, networking, anything that touches the real world has to synchronize with it. For example I was involved with designing a PC graphics chip in the mid '90s - we did the calculations around metastability (we had 3 clock domains and 2 crossings), we calculated that our chip would suffer from metastability (might be as simple as a burble on one frame of a screen, or a complete breakdown) about once every 70 years - we decided we could live with that as they were running on Win95 systems - no one would ever notice Everyone who designs real world systems should be doing that math - more than one clock domain is a no no in life support rated systems - your pacemaker for example
- elwell 13y agoThis field really interests me.
- caf 13y agoIf a failure mode was likely to happen once every 70 chip-years of operation, then it seems like if you sold a few hundred thousand chips then you would expect several instances of that failure mode to occur across the population of chips every day?
- Taniwha 13y agosimply yes - but as I mentioned in our case by far the most most were going to be pixel burbles - you'd likely see one in the lifetime of your video card - the chances of the more serious sort of jabbering core sort of meltdown are much less likely - we design against them - but, one has to stress, not impossible. You can design to be metastablity tolerant - use high-gain, high clk->Q flops as synchronizers, uses multiple synchronizers in a row (trading latency for reliability), you can do things to reduce frequencies (run multiple synchronizers in parallel, synchronize edges rather than absolute values etc), but in the end if you're synchronizing an asynchronous event you can't engineer metastability out of your design - you just have to make it "good enough" for some value of good enough that will keep marketing and legal happy. It's our dirty little secret (by 'our' I mean the whole industry)
- rdtsc 13y agoThere was an interesting quote/anecdote, Joe Armstrong likes to tell, it is about people who claim they've built a reliable or fault tolerant service. They would say "This is fault tolerant, they are multiple hard drives in there, I have done formal verification of my code and so on..." and then someone else trips over the power cord and that's the end of the fault tolerance. It is just a silly example, of course they'd properly provide power to an important rack of hardware, but the point is, in the simplest case the system is only as fault tolerant as its weakest components. It is that one bad capacitor from Taiwan that might the whole thing down, or just a silly cosmic ray. One needs redundant hardware to provide certain guarantees about the service being up. This means load balancers, multiple CPUs running the same code in parallel and comparing results, running on separate power buses, different data centers, different parts of the world.
- shurcooL 13y ago> different parts of the world. Still takes just one asteroid.
- oconnor0 13y agoSeems like after an asteroid hitting the earth, "server fault tolerance" is the least of our worries.
- klodolph 13y agoYeah, but with the number of 9s you see you realize that asteroids are NOT taken into account. For example, Amazon advertises 99.999999999% durability for a given year for S3 objects. This is just stupid. An extinction-level event (asteroid, global thermonuclear war, black hole) could easily wipe out ALL data on S3. We know that mass extinctions have occurred about once every 100 million years. That means that if we expect a 10^-8 chance of a mass extinction event in a given year, Amazon would need a 99% chance of surviving a mass extinction in order to meet average durability ratings for S3. After a certain number of 9s you just have to smile, nod, and truncate the number.
- synthos 13y agoSoft errors are a very real property of low-voltage digital electronics. I personally observed what could only be realistically explained as a soft error in a unit running customer hardware in the field. A single bit was flipped in the program memory of the embedded application and was causing the system to malfunction in an obvious and repeatable manor. We've since added CRC checking to the program memory and some of the static data sections to flag and reset this in the future.
- mvanveen 13y agoMy immediate reaction is to ask how this reliability characteristic of CPUs affects critical software applications? Certainly some space missions and medical devices out in the field must have surpassed the MTBF mark for the given CPU deployment.
- lispython 13y agoThere's a more than 100 pages's thread talk about GUP failure after two years use in Apple Support website. https://discussions.apple.com/thread/4766577 https://discussions.apple.com/thread/4766577
- csmuk 13y agoNever had a CPU go on me. RAM yes, PROMs yes, CMOS batteries yes, PSUs yes, drives yes. They're probably the most reliable bit of a computer.
- mrich 13y agoAs a side note, the whole site is an amazing collection of wisdom and worth bookmarking: http://yarchive.net/ http://yarchive.net/
- jokoon 13y agoI always wondered about this, but does it seem transistor do wear off over time ? Does that mean a CPU/RAM/GPU will not perform as well as when it's brand new ?