5 ms·
All DDR5 has ECC built into the modules for it to reach parity DDR4. ECC memory modules are better because they can fix and report errors, the base DDR5 DIMM w
by undersuit 3y ago
All DDR5 has ECC built into the modules for it to reach parity DDR4.
ECC memory modules are better because they can fix and report errors, the base DDR5 DIMM will hand you a piece of memory with multi-bit corruptions that it couldn't fix. The ECC DIMM will fix larger errors and tell you about each detected corruption.
But there are ECC DIMMs with even better ECC. DDR5 moved the DIMMs from one 64-bit channel to two 32-bit channels. ECC DDR4 you'd protect 8 memory dies with 1 memory die holding the ECC, your computer will identify it as 72-bit memory. ECC DDR5 you protect the 4 memory dies in the half channel with a dedicated 5th die on each channel. Some ECC DDR5 will have 80-bit ECC, some will have 72-bit ECC. You want the 80-bit ECC because it stores 8 more bits of ECC for your memory controller to correct errors.
Micron isn't known for making 80-bit DIMMs.
- jeffbee 3y agoI would add that "has ECC" is even more complicated than you've implied here. There are many ways to detect and correct DRAM errors, and with the memory controllers moved into the CPU the features you get depend on your CPU vendor, firmware, and operating system. To get the most out of a current Xeon SP, for example, you want 10x4 DDR5, that provides 8 code bits with each 32 bits of data. All of the players in the business have patents on some part of ECC, so you get a grab bag of techniques on different platforms.
- crotchfire 3y agoECC is not complicated. Manufacturers pretend it's complicated in order to sell you inferior product at inflated prices. It's really simple: just demand SECDED. That's all you need to know, one acronym, SECDED. https://cr.yp.to/hardware/ecc.html https://cr.yp.to/hardware/ecc.html Single Error Correction, Double Error Detection. BTW, this sort of ECC fraud has been going on since the 1990's at least. It isn't going to just go away. Learn what SECDED is and demand it. On Linux: # cat /sys/devices/system/edac/mc/mc*/*/edac_mode SECDED SECDED SECDED SECDED SECDED SECDED SECDED SECDED
- wtallis 3y agoIt really isn't that simple. For example, the distinction you're ignoring from the comments you replied to is the question of correcting a single error among how many bits: one bit error across the full 64-bit width of a DIMM, or one bit error per each 32-bit sub-channel now that DDR5 has split the DIMM in a way that previous desktop memory standards don't. Where a non-ECC DIMM is usually 8 DRAM chips (64 bits) and in DDR4 or earlier and an ECC DIMM would be 9 DRAM chips (72 bits), now with DDR5 can have either 9 DRAM chips or 10 DRAM chips (80 bits). The question of whether ECC protection is strong enough to provide SECDED also is insufficient to address the complexities of having ECC protection only on the link, or end-to-end via sideband or in-band ECC, or using a combination of separately implemented link ECC and on-die ECC.
- crotchfire 3y agoIt's really that simple. > the complexities of having ECC protection only on the link That's not SECDED, since it can't correct a bitflip that occurs somewhere other than on the link, nor can it detect a double bitflip in those situations. SECDED. Just ask for SECDED. Even the Linux kernel knows about this: # cat /sys/devices/system/edac/mc/mc1/csrow3/edac_mode SECDED SECDED. Just ask for SECDED.
- wtallis 3y ago> That's not SECDED, since it can't correct a bitflip that occurs somewhere other than on the link You're just making up arbitrary rules. The term SECDED does not incorporate any guarantees about which part of a system it applies to. It's just a mathematical statement about the strength of the ECC in use plus the implication that detectable but uncorrectable errors are reported. What you're looking for requires more words to fully specify, probably including the term "end to end".
- jeffbee 3y agoYour mantra is sort of hilarious because with current technology SECDEC is the worst ECC you can buy. If that's what you are getting on a DDR5 server today, the most likely reason is your integrator has made a mistake.