4 ms·
I was under the impression that all DDR5 modules have ECC in them, but only "ECC" modules reported errors to the OS. Wouldn't normal DDR5 be sufficient for your
by hda2 4y ago
I was under the impression that all DDR5 modules have ECC in them, but only "ECC" modules reported errors to the OS. Wouldn't normal DDR5 be sufficient for your workstation?
- Dylan16807 4y agoThe internal ECC is useful but a separate feature. Traditional ECC, achieved by adding extra storage, is still available and will also protect the data as it's sent across the bus. Sadly it now requires 25% extra chips instead of 12.5% extra chips, because of DIMMs being split into two subchannels. Also sadly there's no way to generate a bus-level checksum on the fly. I don't know if internal ECC errors ever get reported.
- lynguist 4y agoSo what happened is the following: When JEDEC was standardizing DDR5 they mandated on-die ECC which checks for errors that happen while the data rests on RAM. This was done to increase reliability because of high memory density. This is good and every DDR5 module has it. However there’s another ECC: the one where errors are checked each time data is transferred between CPU and RAM. This requires the support of the CPU, the RAM, and the mainboard. This is important for data integrity and every computer that runs for multi-day periods must have it. I’ve found by reading papers about it that 1 bit error per 4 GB RAM per 3 days is guaranteed to happen. If RAM is marked DDR5 it has on-die ECC. If RAM is marked DDR5 ECC it has on-die and transfer ECC. That is what I want and it is not optional at all in my opinion.
- DeathArrow 4y ago>I’ve found by reading papers about it that 1 bit error per 4 GB RAM per 3 So for my 64GB of RAM I get 2 bytes worth of errors / day. And a few hundred for 2 or 3 months since I didn't restart my PC.
- bruce511 4y agoThis could easily be true, and easily have zero impact on anything you are doing. Memory can be assigned to all sorts of things so the impact of a single bit change will depend entirely on what that memory is being used for at that moment. For example, I load a program into memory, but I'm not running all of the code in that program all the time. Let's say the program includes code for exporting a list to excel, but that's a feature I never use. Or it has a procedure which uses a local variable as a loop counter. Outside the confines of that loop the variable is meaningless. Plus its unlikely you are even nominally using all 64g at the same time. The errors might all be happening in unused ram. Even if it flips actual data, it may end up being no more than a typo in a document. So yes, a single bit in 4gb might flip from time to time. But the probability of the flip being "meaningful" is less than 1.
- jl6 4y agoIMHO a typo in a document is a serious error that we should put substantial effort into preventing. There are even worse potential outcomes: * A single bit error in a stream of compressed or encrypted data stream that renders the stream unreadable. * A permissions check being yes instead of no. * Backups being silently corrupted
- bruce511 4y agoWhile a typo can be significant, statistically speaking it almost certainly isn't. Equally I'm not saying that all bit flips are inconsequential. I'm saying that most bit flips likely go either completely undetected, or have meaningless consequences. Saying that say 1 bit per 4Gb will flip every day is a long way from saying that a _significant_ (or even detectable) bit flip will happen every day (per 4Gb). From a simple feel for "what memory is used for", coupled with "how much of the ram is actually used" makes me suspect that it's a very low probability of it being an issue. Which makes the parent comment, about 1600 bits of data being flipped since the machine last rebooted [1] being both "true" and likely "meaningless". [1] I'm not sure what rebooting has to do with anything though.
- Teever 4y agoI get what you're saying but on the other hand isn't bit-flipping that results in typos or crashes something that is so beneath 2022? Like isn't this the kind of thing that Turing and Zuse had to deal with? And if it isn't, when is it a problem that we will finally and absolutely excise? Is it more of a 2030 thing? 2040?
- zasdffaa 4y ago> I’ve found by reading papers about it that 1 bit error per 4 GB RAM per 3 days is guaranteed to happen. That is very, very dubious indeed. I can't believe there's no ECC between ram and cpu, it's not credible for anything but a cheap home machine. Any sources?
- Tuna-Fish 4y agoIt sounds wrong because it is wrong. In general, error rates in the modules scale with the physical volume of the chips, not the amount of memory they contain (because the primary root cause is cosmic rays or radioactive contamination inside the chips). Error rates on the bus mostly depend on external circumstances, and scale with transfer amounts, not memory amounts. Most paper characterizing errors are very old (so they used memory chips with large feature sizes), and they scaled their error rates by memory amount (because they didn't know better). This results in massively overinflated error bit rates. It's possible to prove that this is wrong very easily by just writing a program that allocates a large pool of ram, and constantly reads it and writes back to it over a period of time, and then check results after a while. I once did this for an argument on the internet about this -- on my old Intel Sandy bridge system with 16GB non-ecc ram, I allocated a pool of 8GB, wrote ascending integers from 0 to 2^31 to it, and then iterated over it, just reading it in, checking if it's still correct and terminating if it's not, and writing it back out to same address. This program ran for a week with no errors before I had to end it because I needed the machine. 1 bit error per 1.333GB*days is completely not credible.
- temac 4y agoI don't completely understand your conversation. Maybe the error rate is inflated, but as for: > . I can't believe there's no ECC between ram and cpu, it's not credible for anything but a cheap home machine. it is absolutely not wrong: there is no ECC between CPU and RAM for most consumer and desktop computers, and even tons of embedded systems, and that's basically the only bus lacking even basic integrity checks, so there really should be some ECC also there, nearly everywhere.
- 4y ago
- PragmaticPulp 4y ago> I’ve found by reading papers about it that 1 bit error per 4 GB RAM per 3 days is guaranteed to happen. This estimate is several orders of magnitude too high. We’d be seeing weird bit flips and file corruption all over the place if this was even close to being true. I have ECC systems that will report the number of errors detected and corrected. Still waiting to see any errors on my main machine with 64GB of RAM. At server farm scale, the majority of memory errors come from a very small number of faulty memory sticks. It’s not an even distribution of errors across all memory.