3 ms·
> Ok so tell me what it looks like. _That's_ what I want to know. Suppose an application running on your machine suffers some type of malfunction -- a segmenta
by theevilsharpie 3y ago
> Ok so tell me what it looks like. _That's_ what I want to know.
Suppose an application running on your machine suffers some type of malfunction -- a segmentation fault or a seemingly-random kernel panic that you're unable to reproduce.[1] And suppose it happens often enough that you want to fix it, and understand the root cause.
You research the issue, and are pointed to potentially faulty hardware. That then raises a question: how do you know your physical memory is working properly?
You could run a diagnostic application like Memtest86. However, that has the following issues:
- It only tests a particular point in time. If the test "passes", that doesn't say whether your memory encountered a fault in the past, nor does it guarantee that it won't fault in the future.
- What does it even mean for a test to "pass?" Just a single run through a test suite? Running it for some period of time, like a few hours?
- A diagnostic like Memtest86 is invasive. While it's running, you can't use your machine for other things.
Troubleshooting these types of hardware errors without ECC is tricky, because there's not really a way to conclusively link a software fault to a hardware error.[2][3]
However, on a machine equipped with ECC, this type of troubleshooting is a lot more straight-forward. If bits are flipped in memory, the memory controller can probably detect that (assuming only only one or two bits are flipped), and can raise a machine exception that the OS can catch and do something with (e.g., logging the error, terminating the process impacted if the error is correctable, etc). That saves a lot of time and headache that you'd otherwise be spending on guesswork.
You may still be asking, "how big of an issue is this, really?" I suppose whether or not you care about having hardware that can detect memory faults is up to you.[4] However, during the portion of my career where I was managing large machine fleets, memory failures were the second most common type of hardware failure, behind only mechanical disk drives.
> That statement doesn't hold water. If it takes extra bandwidth and capacity to provide ECC, I can use the extra bandwidth as memory instead of error correction, no?
No. The memory controller is unable to use the additional capacity for anything other than error checking (hence why it's called "side-band").
What you're describing is "in-band ECC", which is something that you sometimes see on GPUs or small form factor systems aimed at professional markets.
---
[1] An application crash is actually one of the better outcomes of a memory error. The worst-case scenario is silent corruption of data stored in memory that your application assumes to be valid.
[2] Errors as a result of faulty memory have been particularly frustrating to OS developers, as it can result in nonsense error reports. Linus Torvalds has bemoaned the lack of ECC options on consumer hardware, claiming that it was the industry cheaping out. During the lead-up to the launch of Windows Vista, Microsoft officially encouraged the use of ECC. I don't follow development of the BSDs, but it wouldn't surprise me if they similarly wished their users would use ECC across the board.
[3] When Google first started building out the infrastructure for their server farms, they opted not to use ECC memory for their servers as a cost-savings measure. They subsequently had a memory fault on one their machines that resulted on corruption of their search index. While they worked around the problem by adding logic to verify the contents of the index, all subsequent generations of machines that Google has deployed have ECC support. For specifics, see the second footnote in https://danluu.com/why-ecc/ https://danluu.com/why-ecc/
[4] Note that many of your other components that store data in some way (e.g., caches, persistent storage, etc.) are likely to have ECC capability. The only notable exceptions I can think of are GPU memory, the main system memory on consumer hardware, and the processor's registers.