3 ms·
Is there any windows / linux utility that shows the number of ECC errors corrected since the boot? If not, AMD should write and promote one. Very curious on h
by srcmap 9y ago
Is there any windows / linux utility that shows the number of ECC errors corrected since the boot?
If not, AMD should write and promote one.
Very curious on how often it happened for normal home/offic usage.
I used to work for a silicon company who took a embedded network switch system with ECC logic to some nuclear lab for testing to verify/showcase the ECC functionalities.
- mjevans 9y agoEven if you never personally encounter an error, knowing that the 1 bit correct and 1+ bit likely detect exists will save you in terms of piece of mind and trouble-shooting other issues. Plus, if you're scrubbing your storage the last thing you want is a memory error killing your data.
- planteen 9y agoIn Linux, yes, there is a service called mcelog and a utility from the edac-utils package called edac-util. You will see correctable ECC errors on systems. How frequent honestly seems to depend on the workload and the system itself. My suspicion is that often they are caused by poor PCB layout and ECC saves you. I spent literally weeks (nights, weekends) chasing down an issue I thought was a software bug but turned out a board layout issue on an embedded system. If the system had ECC, the error would have either been corrected or we would have gotten the uncorrectable ECC error trap. Since then, every workstation/server/desktop I spec is ECC. I wish more laptops had it.
- srcmap 9y agoThanks, Try the edac-util on 5 of 30+ or so servers. "Intel(R) Xeon(R) CPU E5-2667 v3 @ 3.20GHz" with 128G RAM each. edac-util: No errors to report. System uptime ~30+ days. We recently have to move those servers. I wish I check this command before the move. The uptime should be 600 days + for some of the servers.
- imtringued 9y agoThe problem with bit flips is that they accumulate in high uptime systems. If you reboot your PC at least once every week it's not going to be a problem.
- JoeAltmaier 9y agoEarly ECC days, you 'washed' memory to fix this. On a read, a single-bit ECC error is actually repaired by the hardware. To get the most benefit from this you would want to read every allocated memory location periodically, 'washing it clean' so the accumulated errors wouldn't become double-bit errors (unrecoverable). I'd put a wash routine in the background process, where it would string-move a block of memory to nowhere in a round-robin way. Not a terrible hit on the cache; we're idle when in the background task so not impacting the most used code. Some latency issue with interrupts and the like.