5 ms·
Maybe slightly OT, since this concerns AMD older AM4 platform with a Zen3 APU core, but working ECC support looks like this and is definitely present on my syst
by c0l0 3y ago
Maybe slightly OT, since this concerns AMD older AM4 platform with a Zen3 APU core, but working ECC support looks like this and is definitely present on my system:
$ sudo ras-mc-ctl --errors | tail -n5
14 2023-08-20 20:16:41 +0200 error: Corrected error, no action required., CPU 2, bank Unified Memory Controller (bank=17), mcg mcgstatus=0, mci CECC, memory_channel=0,csrow=0, mcgcap=0x0000011c, status=0x9c2040000000011b, addr=0x36e701dc0, misc=0xd01a000101000000, walltime=0x64e31c78, cpuid=0x00a50f00, bank=0x00000011
15 2023-08-23 17:17:49 +0200 error: Corrected error, no action required., CPU 2, bank Unified Memory Controller (bank=17), mcg mcgstatus=0, mci CECC, memory_channel=0,csrow=0, mcgcap=0x0000011c, status=0x9c2040000000011b, addr=0x36e701dc0, misc=0xd01a000101000000, walltime=0x64ea5188, cpuid=0x00a50f00, bank=0x00000011
16 2023-09-03 16:52:15 +0200 error: Corrected error, no action required., CPU 2, bank Unified Memory Controller (bank=17), mcg mcgstatus=0, mci CECC, memory_channel=0,csrow=0, mcgcap=0x0000011c, status=0x9c2040000000011b, addr=0x36e701dc0, misc=0xd01a000101000000, walltime=0x64f4d227, cpuid=0x00a50f00, bank=0x00000011
17 2023-09-15 21:37:59 +0200 error: Corrected error, no action required., CPU 2, bank Unified Memory Controller (bank=17), mcg mcgstatus=0, mci CECC, memory_channel=0,csrow=0, mcgcap=0x0000011c, status=0x9c2040000000011b, addr=0x36e701dc0, misc=0xd01a000101000000, walltime=0x65071ed7, cpuid=0x00a50f00, bank=0x00000011
This is with an ASRock B550M-ITX/ac and a AMD Ryzen 5 PRO 5650G. It used to work the same with a Ryzen 5 3600 (using a dedicated GPU for video output) before I upgraded the CPU.
To detect and log ECC activity on modern GNU/Linux, you will want to have the "rasdaemon" service active. I will decode MCE (and other hardware-related) errors and persist them to the database that is shown being queried above.
- jacquesm 3y agoThe frequency of these errors should be enough to cause you to distrust any kind of information output by a computer without ECC. Edit: thinking about this a bit longer: that frequency is actually so high that you may well have a broken module in there. Note how it is the same module and the same address every time.
- code_biologist 3y agoIt'd be cute to log what processes are using that memory at time of error. Fun to speculate about whether a kernel bit flip is better or worse than ones in a web browser, photo editor, spreadsheet, network storage client...
- jacquesm 3y agoYou could test that empirically by setting up a box with the express intent to crash it and then using a chaos monkey like mechanism where you start injecting single bit faults into memory at random addresses. Wonder how long the box would be up before you start noticing something is broken. It would be funny if you accidentally killed the chaos monkey first! Best not use that box for banking...
- whizzter 3y agoSeems like a semi-stuck bit, it'd definitely cause issues w/o ECC but seems to chug along with the circuits doing their job. Best would probably be to add an memory-range exclusion to the kernel at boot to avoid that single area since the sticks seems good otherwise.
- Modified3019 3y agoI completely forgot you could do that.
- nvarsj 3y agoIme ram is either bad or good. I’ve had ecc errors like this and I always ask the DC to replace the ram. After that, 0 errors forever. Same reason why I’m confident a 24 hour memcheck is sufficient for non ECC ram.
- simcop2387 3y agoGenerqlly the same here, but I have had sticks fail after some time in use. I had to rma the ram in my frame.work laptop after it failed. No reason or clue why but it happened after 6 minths or so. No issues with the rma though and it went fine with no issues since. ecc if it was supported there might have given me a heads up about it and avoided needing me to restore from backups when the fs corrupted.
- c0l0 3y agoI am fully aware the module is not 100% working, i.e., it is faulty at a specific physical address. That's OK for my personal desktop though, unless the condition worsens, and UCEs (which will panic my kernel) follow.
- justinclift 3y agoAny chance it's something that's being affected by temperature? Along the lines of "computer gets toasty doing work, ecc errors start happening"? Stuff like that could just mean the memory sticks need pushing in a bit more.
- c0l0 3y agoI tried to test and control for that to the best of my ability, but ambient/operating temperature does not seem to be part of the equation.
- _cenw 3y agoAPUs are specifically excluded from supporting ECC, except on the PRO SKU. https://www.asus.com/global/support/FAQ/1045186/ https://www.asus.com/global/support/FAQ/1045186/
- just_testing 3y agoWhich memory sticks do you use?
- c0l0 3y agoTwo sticks of Kingston 9965745-042.A01G
- just_testing 3y agoThanks!
- cherryteastain 3y agoI have ECC RAM installed on my Gigabyte B550I system. dmidecode shows the 72 bit width (Total Width: 72 bits) and dmesg | grep -i EDAC does show a bunch of info suggesting ECC is enabled. But this command's output is empty: No Memory errors. No PCIe AER errors. No Extlog errors. No MCE errors. Do I need to enable something so these errors get logged or have I been misled by dmidecode and dmesg?
- c0l0 3y agoYou are just lucky, and your hardware appears to be working without any problems ;) Some/most(?) AM4 boards can enable "PCIe AER" (Advanced Error Reporting) in their firmware, which will tell you about stuff going awry while components are communicating over said bus (but every instance of PCIe error I have ever seen, even on rather faulty hardware, was recoverable/correctable), and rasdaemon will also persist those. I do not know what "Memory errors" are supposed to be, since ECC-related problems will be dropped into the "MCE errors" bucket. Neither do I know what "Extlog errors" are.