3 ms·
I ran memtest+ for a weekend every 3 months or so on 64GB DDR4 Ryzen system for 2+ years and haven't found anything. Rock solid system, which has been on almost
by ece 6y ago
I ran memtest+ for a weekend every 3 months or so on 64GB DDR4 Ryzen system for 2+ years and haven't found anything. Rock solid system, which has been on almost 24/7. No data loss that I can verify with backups.
I think the point about chipkill is salient, why isn't it more widely used if the benefits there are obvious? I think ECC is completely not worth the price for home users if you can do regular backups, run memtest, and compare checksums for any data corruption. Even with errors near the hard/soft boundary, you'll likely catch those by just running memtest for a longer period of time. DRAM errors get progressively worse, so you'll catch any errors that way as well using memtest again.
Datacenters have more surface area which cosmic rays can attack and are also more likely to see weird hard errors which might not have been caught in QC/QA, which is very different from isolated home computers or phones which aren't on 24/7 or can restart easily. If you have a datacenter and have truly sensitive unreplicated data, get chipkill. If you are home user, do regular backups, and run memtest. It's that simple. Hard drives have moving parts, and more failure modes, so pick a file system with checksums. CPUs/GPUs and SSDs might or might not have ECC caches because they have other ways of reducing hard/soft errors in SRAM.
People may disagree, but the study I would like to see is one that takes into account the denser environment of a datacenter vs the isolated home user and all modes of data loss/corruption in each case. We know Google replicates data in GFS, and a home user can do the same with backups.
- account42 6y ago> I ran memtest+ for a weekend every 3 months or so on 64GB DDR4 Ryzen system for 2+ years and haven't found anything. Don't assume that that means there are never any errors. DRAM errors often depend on access patterns: While dialing in the speed and timings for my RAM initially settled on a configuration that did not report any errors in memtest but later during heavy usage (e.g. compiling LLVM with make -j64) reported ECC errors.
- ece 6y agomemtest exactly tests for different access patterns including attacks like Rowhammer. I also use Gentoo with background tasks like mprime/foldingathome (they can catch memory errors too), and any gcc/llvm errors would've been obvious by now (It's easy to repeatably and reliably test checksums on every compiled package). DRAM errors get progressively worse, replace the module when you find any errors. Simple as that. Focusing on finding errors early and doing backups protects against almost all classes of failures for home PCs. With a median of 10,000 FIT or lower, it's zero errors every 5-10 years, and DDR4/5 might actually be below that if the rate of improvement between DDR1/2 was anywhere between 2x-10x continued. You'd have to overlook this important trend to see any benefit from ECC. Yes, there are DRAM errors, but if they can be found by memtest, bg scrubbing, self-tests and don't affect low-use scenarios, then it's meaningless to be talking about ECC for home use. Setting speed and timings on ECC modules seems super flaky to me, do you know what speed and timings the separate ECC logic can handle? Can you turn off ECC and test DRAM independently? Maybe what you're seeing is the ECC logic throwing errors and not the DRAM.