3 ms·
This happened before a couple of years ago, and the problem seems to have repeated itself. https://news.ycombinator.com/item?id=27728287 https://news.ycombinato
by nickf 3y ago
This happened before a couple of years ago, and the problem seems to have repeated itself. https://news.ycombinator.com/item?id=27728287 https://news.ycombinator.com/item?id=27728287
- topspin 3y agoIs coincidence plausible here? Anyhow, Ian Fleming's thoughts on such things seem applicable: "Once is happenstance. Twice is coincidence. Three times is enemy action"
- jeroenhd 3y agoBitflips just happen. There's a good chance the device and connection you're using right now has had several bit flips in (unused) memory and you'll never even know. Everything is fine until these flips happen in critical code or data paths. ZFS corruption is a famous example, but there's also DNS corruption that happens quite regularly (and has been demonstrated to be usable for malicious purposes). I'm a little surprised the certificate transparency protocol doesn't validate the incoming data well enough to detect these bit flips, but on the other hand most software I've seen just assumes the bytes in memory and the bytes received through the network are all what they're supposed to be.
- theamk 3y agoeh, I bet you are wrong. I've run memtest many times, often for days, and it never detected any bitflips in RAM. And memtest is specifically designed to exercise every bit of memory and detect bitflips. While I've seen my share of PCs with memory errors, when you see one you replace memory / tweak settings, and then run memtest for a day or so to ensure it does not reappear. Also, "unused memory" should not be a thing in modern PC: the read cache should expand to fill it - it's super fast to evict and provides tangible benefits in case of hit. The errors usually occur on the boundaries - like in the network or in the SATA connection to disk or even in USB bus. For example the original "ZFS corruption" story (which seems t be gone from regular web, but I think its [0]) pretty clearly mentions damage "on the way to disk". [0] https://web.archive.org/web/20091212132248/http://blogs.sun.com/elowe/entry/zfs_saves_the_day_ta https://web.archive.org/web/20091212132248/http://blogs.sun....
- sliken 3y agoNot sure what it is about memtest86, but in my experience it doesn't find most memory errors. I've had serious memory issues, that trigger memlogd, dmesg, or similar about EDAC/ECC errors. I try memtest86 over night, no errors. Restart the node and get more errors. This seems to happen most of the time, only in the rare case can I reproduce errors I see in a production system with memtest86 reports.
- ThePowerOfFuet 3y agoSame guy too (Andrew Ayer). Someone is keeping an eye on it!
- KirillPanov 3y agoI cannot understand why they don't use two-of-three voting to produce the log entries. It doesn't have to be three machines owned by separate organizations, or even in separate buildings. Just three servers in the same rack, doing the same computations, and nobody signs anything unless one of the other two produces the exact same result.