4 ms·
Something important not discussed here is that you need not only error correction (when possible) but also error reporting, so you can know you RAM is failing.
by nahuel0x 2mo ago
Something important not discussed here is that you need not only error correction (when possible) but also error reporting, so you can know you RAM is failing. Sadly, is not possible to get DDR5 on-die ECC error counts reported to the OS. Alternative solution, use IBECC with DDR5, where the CPU reserves a portion of the RAM for parity bits and does the check / reporting.
- inigyou 2mo agoMy Linux installation came with an AMD MCE driver that reports them to dmesg. It logs a few corrected errors a day.
- rubatuga 2mo agoIsn't that a lot? I see them only when heavily over clocking .
- adrian_b 2mo agoA few errors per day seems much too frequent, unless you live at a high altitude. Normal good DIMMs, at least when new, should not have errors more frequently than one error per many months. Frequent errors may appear with memories that are not seated well in their sockets, or which are old, at least several years old. Frequent errors may also be caused by more general computer problems, like a bad power supply unit.
- bikelang 2mo agoCurious to learn more here - why would altitude play a role here? And when you say high altitude - are we talking La Paz, Bolivia (~12k feet) or Denver, CO (~5k feet)?
- dsr_ 2mo agoThe primary source of uncorrelated bit-flips is, effectively, cosmic radiation. (Correlated bit-flips are likely from a manufacturing error... or, at least in one case, from excessively radioactive ceramic packaging.) The more atmosphere you have to randomly absorb high-energy photons, the better. Higher density memory drops fewer electrons in each well, which means a lower-energy photon can change the state. Corolary: submarine datacenters are much better than orbital datacenters. Cheaper to cool, construct, shield and access.
- willis936 2mo ago>The primary source of uncorrelated bit-flips is, effectively, cosmic radiation. This gets said a lot but with little evidence. The mario speedrun bit flip has been replicated with marginal connector insertion. https://youtu.be/vj8DzA9y8ls https://youtu.be/vj8DzA9y8ls
- adrian_b 2mo agoThere is a lot of direct evidence from increased error rates in computers on airplanes.
- willis936 2mo agoThat's a much weaker statement than "the primary source of uncorrelated bit flips is cosmic radiation".
- adrian_b 2mo agoCosmic radiation is always a source of errors for any kind of DRAM. Whether it is the primary source of errors, depends on the design of the DRAM die. In a badly designed DRAM there could be many other error causes. The designers of a DRAM certainly attempt to minimize all the error causes that they can control. If they succeed, the cosmic radiation would remain the primary error source. Without access to internal documents of the DRAM vendors, we cannot know if they have indeed minimized all other error causes. In modern DRAM chips, a major error cause can also be the disturbance of some memory cells when some other cells are accessed in their neighborhood, which is the basis of the Rowhammer attacks, but which can also happen during the normal use of the memory, when certain access patterns happen by chance. The DRAM vendors do not provide details about the behavior of their devices, with the hope that this makes harder the task for someone who wants to design a variant of the Rowhammer attacks, but this also makes impossible for the owner of the memory to predict whether an application program will not perform by chance an access pattern that will unintentionally flip some memory bits, which without ECC will not be detected and corrected.
- fvwqcecvq 2mo agoAre DIMMs (a.k.a. UDIMM) still a thing? I would think most servers and workstations would be RDIMM (Registered DIMM) by now and consumer stuff uses soldered down memory. Memory failing because it's old is definitely a thing, and very possible in this scenario, but I feel like I haven't seen errors due to physical insertion, that were not caught immediately by POST, in years. But maybe it's just me. Happy to be lucky. :)
- wtallis 2mo agoThe client/consumer desktop market hasn't actually disappeared, and soldered memory is almost unheard-of in that market segment.
- etbe 2mo agoDoes this imply that the all plastic gamer PCs (that are more designed to be art projects than serious computers) are bad and that everyone should be using systems with solid steel cases to block radiation?
- adrian_b 2mo agoA steel case has too little effect on cosmic radiation, but it is good to shield electrical noise. The computers in plastic cases can sometimes be affected by appliances with electrical motors that are used nearby them. The cosmic radiation is attenuated only inside a big building, e.g. one with many stories above you. A really great attenuation is obtained only underground.
- etbe 2mo agoGood point. Some googling suggests that plastic may be better than metal as hydrogen absorbs high energy protons without generating secondary radiation.
- scheme271 2mo agoPlastic cases might actually be a bit better than steel depending on the radiation type. Neutron radiation is more effectively shielded by interacting with light atoms so stuff with lots of hydrogen atoms like water or polymers have a better chance of interacting with neutrons and slowing it down so that it can be stopped. Steel is probably better for gamma radiation or charged particles like muons. But overall, pc cases are probably too thin to appreciably affect the radiation levels. Cosmic rays and their secondary radiation is fine with going through buildings so a 1mm thick case isn't going to do much.
- loeg 2mo agoCheck if those are all coming from DIMMs, as opposed to other subsystems (on-die caches).
- etbe 2mo agohttps://www.theregister.com/on-prem/2021/06/04/fyi-todays-computer-chips-are-so-advanced-they-are-more-mercurial-than-precise-and-heres-the-proof/598198 https://www.theregister.com/on-prem/2021/06/04/fyi-todays-co... Google and Facebook report a few unreliable CPU cores out of thousands of systems. Meaning maybe a 1/10,000 rate of repeatable CPU errors. https://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf https://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf Google reports 8% of DIMMs being affected by errors. While the differences in metrics prevent comparing the two things directly it seems that DRAM errors are more common than CPU errors.
- sliken 2mo agoSounds like an issue with hardware. I watch ECC errors for 2500 servers and see a few some days, and zero others. Said servers at an elevation of 6,000 feet or so. A few per day for a single machine is crazy high. You likely have a dimm, dimm slot, motherboard, or CPU issue that's susceptible to RF noise, vibration, temp changes or similar.