8 ms·
Might be a noob question but for truly important data, couldn't SDCs be detected by using ECC everywhere?
by opisthenar84 3y ago
Might be a noob question but for truly important data, couldn't SDCs be detected by using ECC everywhere?
- jandrewrogers 3y agoECC isn’t free and ECC has a limited ability to detect all statistically plausible errors. Additionally, error correction in hardware is frequently defined by standards, some of which have backward compatibility requirements that go back decades. This is why, for example, reliable software often uses (quasi-)cryptographic checksums at all I/O boundaries. There is error correction in the hardware but in some parts of the silicon that error correction is weak enough that it is likely to eventually deliver a false negative in large scale systems. None of this is free, and there are both hardware and software solutions for mitigating various categories of risk. It is explicitly modeled as an economics problem i.e. how does the cost of not mitigating a risk, if it materializes, compare to the cost of minimizing or eliminating it. In many cases, the optimal solution is unintuitive, such as computing everything twice or thrice and comparing the results rather than using error correction.
- lobochrome 3y agoIn those cases, the CPU makes a false calculation independent of what's done in RAM. It can be solved by having flop redundancy as in system z - but nobody at Google or Meta would be considering big metal. From my point of view, this technology problem may be interesting academically (and good for pretending to be important in the hierarchy at those companies) but a non-issue at scale business-wise in modern data centers. Have a blade that once in a while acts funny? Trash and replace. Who cares what particular hiccup the CPU had.
- delroth 3y ago> a non-issue at scale business-wise in modern data centers. I've worked on similar stuff in the past at Google and you couldn't be more wrong. For example, if your CPU screwed up an AES calculation involved in wrapping an encryption key, you might end up with fairly large amounts of data that can't be decrypted anymore. Sometimes the failures are symmetric enough that the same machine might be able to decrypt the data it corrupted, which means a single machine might not be able to easily detect such problems. We used to run extensive crypto self testing as part of the initialization of our KMS service for that reason.
- lobochrome 3y agoSure. It’s a cool issue to work on and maybe actually relevant at Google scale. But I’ve asked your colleagues multiple time if the business side actually cared about the issue and they never confirmed. Again, cool to work on at Google. Not sure anybody else cares. If you care (finance) you fix it in hardware (system z).
- withinboredom 3y agoWhy would the business side ever care about technical details? It's like asking the business what days the dumpsters get emptied. Nobody gives a fuck; they just care that it gets done and gets done quickly, correctly, and safely.
- lobochrome 3y agoA CFO knows which factors have a significant impact on the bottom line.
- withinboredom 3y agoIf a CFO knows which days the dumpster is emptied, you have a strange CFO. The metaphor is to point out that there’s a lot of technical details that aren’t tracked (like usually refactoring isn’t tracked independently) and shouldn’t be tracked because they are the normal part of the technical job. A CFO can’t measure it even if they wanted to because nobody else is measuring crazy things, like how fast you walk to the bathroom or any other metrics that are specifically related to doing your job.
- teaearlgraycold 3y agoThere are errors within the CPU. As for adding ECC within the CPU, I think that would require you to essentially have a second CPU in parallel to compare against.
- moonchild 3y agoCaches, register files, and coherency traffic all definitely include error-correction.
- XorNot 3y agoYou actually need 3 - which is how it's done for space (I believe SpaceX uses this as a solution to avoiding radiation hardened costs). 2 will tell you if they diverge, but you lose both if they do. 3 let's you retain 2 in operation if one does diverge.
- jorticka 3y agoIf you're not hard realtime 2 is enough, you just redo the computation.
- MertsA 3y agoBut if it's a consistent fault, like the silent data corruption covered in the linked paper, redoing the computation is still going to end up with no way to identify which core is faulty. If it's an intermittent fault, then even for hard realtime you can accomplish that with one core, just compute 3x and go with majority result.
- jorticka 3y agoIf it's consistent and persistent, wouldn't that classify as broken hardware requiring device change? Even with 3 chips, if one is permanently wrong you are then left with only 2 working ones so no redundancy is left for further degradation. > just compute 3x That might be difficult if CPU is broken. How are you sure you actually computed 3 times if you can't trust the logic.