5 ms·
I just want to be clear that in the original referenced article, I am not anti-ECC per se, I just found myself caught in the massive cognitive dissonance betwee
by codinghorror 11y ago
I just want to be clear that in the original referenced article, I am not anti-ECC per se, I just found myself caught in the massive cognitive dissonance between "you must have ECC in all your computers otherwise they will constantly and silently corrupt your data + crash" and "statistically speaking, most computers in the world do not use ECC". How can both of these things be true?
The argument for ECC is credible (I personally think rowhammer is the best example of this actually mattering, but ironically a) you can rowhammer ECC memory just fine and b) DDR4 has hardware features to mitigate rowhammer -- which shows how quickly things are changing), but it also seems to hinge on whether you have hundreds to thousands of computers all working together, e.g. the positive effects of ECC only seem to matter enough statistically at a _very_ large scale.
- rwmj 11y agoThe answer is unfortunately: because we don't act like professionals. I hope that, say, civil engineers wouldn't build bridges the way we throw together computer software and hardware.
- pkaye 11y agoWell they do... just look at the fiasco with the replacement for the Bay Bridge in California.
- scurvy 11y agoThat decision was not made by bridge engineers. It was made by architects, artists, and civilians. Basically, everyone with zero engineering experience. The whole thing was a political rig job, setup to create something iconic, rather than something that worked. San Franciscans insist on always doing things differently, not necessarily doing them well.
- pkaye 11y agoWell lots of people come into the picture when you build any engineering system. This includes software and hardware. There are many cost tradeoffs and compromises.
- scurvy 11y agoIn this case, it was purely non technical reasons and not many tradeoffs. It was all political: http://www.sfchronicle.com/bayarea/article/Bay-Bridge-s-troubles-How-a-landmark-became-a-6021955.php?t=b20c182db7 http://www.sfchronicle.com/bayarea/article/Bay-Bridge-s-trou...
- hannob 11y agoIs there a real-world practical rowhammer attack against ECC memory? I haven't seen one and the test tools I ran never triggered anything on servers with ECC.
- codinghorror 11y agoIf you read the comments to the referenced article, someone had a rowhammer problem even with ECC memory. ECC does indeed reduce the chance of a memory error a lot, but the chance is far from zero (at scale of hundreds/thousands of servers), even with ECC. Another paradox about ECC, it's not a guarantee, so you still need to build systems that can tolerate / mitigate statistically rare memory error states. This commenter seemed quite credible to me, here's where it starts: http://discourse.codinghorror.com/t/to-ecc-or-not-to-ecc/3771/9?u=codinghorror http://discourse.codinghorror.com/t/to-ecc-or-not-to-ecc/377... > Yep, these are what we had -- uncorrectable errors with ECC memory caused by row hammer. Luckily there are mitigations. Sandy Bridge allows you to double...
- Dylan16807 11y agoRowhammer might get through eventually, but it will blow the corrected error count through the roof on the way there. The job of ECC is to prevent transient failures and to warn you about less-transient failures. It's not a paradox that it can't paper over things that are actually broken.
- mikeash 11y agoUncorrectable errors with ECC memory caused by row hammer is just ECC working as designed. The scary thing about rowhammer isn't just the potential a fault, it's the potential for a security vulnerability. Yeah, a DOS attack isn't nice either, but it's a million times less worse than a privilege escalation caused by memory corruption. For rowhammer to be even the same magnitude of problem with ECC memory would require not an uncorrectable error, but an error that gets past the error detection mechanism entirely and produces data either seen as correct or correctable, but which is not the original data.
- slavik81 11y ago> "you must have ECC in all your computers otherwise they will constantly and silently corrupt your data + crash" and "statistically speaking, most computers in the world do not use ECC". How can both of these things be true? What makes you say that? Both those things being true simply means that most computers silently corrupt your data and crash. That matches my experience. My programs occasionally crash. My pictures and videos are occasionally corrupted. Do I know that those events are caused by memory failures? No. Most of them are probably other sorts of software or hardware failures, but some could be memory errors.
- codinghorror 11y agoI've never had a photo or video occasionally corrupted on any computer I've ever owned going back to 1985. Crashes? Sure, who hasn't. "Some could be.." is computing by coincidence, and I'm not a fan of that logic. You can refer to the 2007, 2009, and 2012 studies for measured data on server farms.
- darkmighty 11y agoA few bit flips in an image are unlikely to be noticeable (it depends on the format though; uncompressed is obviously the most resilient, JPEG should be reasonably resistant too. Also, I don't think images stay for too long in RAM. Your HDD definitively has ECC.
- Moru 11y agoActually jpg is not very resilient, there is artforms doing single-bit changes to jpg-files to end up with very strange images.
- darkmighty 11y agoThis is an adversarial error, I think it's pretty good against random errors. I'll do a test sometime.
- Freaky 11y ago
- Freaky 11y ago> "statistically speaking, most computers in the world do not use ECC". How can both of these things be true? ... most computers aren't particularly reliable. I'm curious - are you not monitoring or at least keeping a vague eye on ECC correction events with your existing hardware? If so, are you just not seeing any? I've never really operated at any sort of "large" scale - a handful of racks at most - but I've always found correction events to be about as routine as IO errors, and certainly way more common than outright disk failures.
- codinghorror 11y agoNeither of our (very good IMO) sysadmins bothered to check ECC counters on our existing server memory as far as I know. We did check SMART counters and SSD wear levels, though.. It is certainly a very good point that unusual numbers of ECC errors (corrected) are a bad sign and that machine should probably have its memory replaced at a minimum.
- Freaky 11y agoI.. would probably have done that before concluding I didn't have any need for it. Even just do a quick one-off mcelog run to see what events have happened recently - you don't have to do some big statistical analysis. I'd estimate about half the machines I've used with ECC have had at least one correction a year, maybe a quarter had one every few months. I've seen one-off weird bursts that never happen again, and I've had quite a few cases where ongoing corrections have indicated a DIMM needed reseating. An actual blatantly faulty DIMM's been quite rare.
- cromulent 11y agoYou aren't concerned about cosmic rays? https://en.wikipedia.org/wiki/Cosmic_ray#Effect_on_electronics https://en.wikipedia.org/wiki/Cosmic_ray#Effect_on_electroni... >> Studies by IBM in the 1990s suggest that computers typically experience about one cosmic-ray-induced error per 256 megabytes of RAM per month.[76] To alleviate this problem, the Intel Corporation has proposed a cosmic ray detector that could be integrated into future high-density microprocessors, allowing the processor to repeat the last command following a cosmic-ray event.[77]