4 ms·
Neither of our (very good IMO) sysadmins bothered to check ECC counters on our existing server memory as far as I know. We did check SMART counters and SSD wear
by codinghorror 11y ago
Neither of our (very good IMO) sysadmins bothered to check ECC counters on our existing server memory as far as I know. We did check SMART counters and SSD wear levels, though.. It is certainly a very good point that unusual numbers of ECC errors (corrected) are a bad sign and that machine should probably have its memory replaced at a minimum.
- Freaky 11y agoI.. would probably have done that before concluding I didn't have any need for it. Even just do a quick one-off mcelog run to see what events have happened recently - you don't have to do some big statistical analysis. I'd estimate about half the machines I've used with ECC have had at least one correction a year, maybe a quarter had one every few months. I've seen one-off weird bursts that never happen again, and I've had quite a few cases where ongoing corrections have indicated a DIMM needed reseating. An actual blatantly faulty DIMM's been quite rare.