4 ms·
ZFS detects corruption. A very long ago someone named cyberjock was a prolific and opinionated proponent of ZFS, who wrote many things about ZFS during a time
by Modified3019 1y ago
ZFS detects corruption.
A very long ago someone named cyberjock was a prolific and opinionated proponent of ZFS, who wrote many things about ZFS during a time when the hobbyist community was tiny and not very familiar with how to use it and how it worked. Unfortunately, some of their most misguided and/or outdated thoughts still haunt modern consciousness like an egregore.
What you are probably thinking of is the proposed doomsday scenario where bad ram could theoretically kill a ZFS pool during a scrub.
This article does a good job of explaining how that might happen, and why being concerned about it is tilting at windmills: https://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-your-data/ https://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-y...
I have never once heard of this happening in real life.
Hell, I’ve never even had bad ram. I have had bad sata/sas cables, and a bad disk though. ZFS faithfully informed me there was a problem, which no other file system would have done. I’ve seen other people that start getting corruption when sata/sas controllers go bad or overheat, which again is detected by ZFS.
What actually destroys pools is user error, followed very distantly by plain old fashioned ZFS bugs that someone with an unlucky edge case ran into.
- tmoertel 1y ago> Hell, I’ve never even had bad ram. To what degree can you separate this claim from "I've never noticed RAM failures"?
- wtallis 1y agoIt isn't hard to run memtest on all your computers, and that will catch the kind of bad RAM that the aforementioned doomsday scenario requires.
- Modified3019 1y agoYou can take that as meaning “I’ve never had a noticed issue that was detected by extensive ram testing, or solved by replacing ram”. I got into overclocking both regular and ECC DDR4 ram for a while when AMD’s 1st gen ryzen stuff came out, thanks to asrock’s x399 motherboard which unofficially supporting ECC, allowing both it’s function and reporting of errors (produced when overlocking) Based on my own testing and issues seen from others, regular memory has quite a bit of leeway before it becomes unstable, and memory that’s generating errors tends to constantly crash the system, or do so under certain workloads. Of course, without ECC you can’t prove every single operation has been fault free, but as some point you call it close enough. I am of the opinion that ECC memory is the best memory to overclock, precisely because you can prove stability simply by using the system. All that said, as things become smaller with tighter specifications to squeeze out faster performance, I do grow more leery of intermittent single errors that occur on the order of weeks or months in newer generations of hardware. I was once able to overclock my memory to the edge of what I thought was stability as it passed all tests for days, but about every month or two there’d be a few corrected errors show up in my logs. Typically, any sort of stability is caught by manual tests within minutes or the hour.
- phatskat 1y agoMy friends and I spent a lot of our middle and high school days building computers from whatever parts we could find, and went through a lot of sourcing components everywhere from salvaged throwaways to local computer shops, when those were a thing. We hit our fair share of bad RAM, and by that I mean a handful of sticks at best.
- wtallis 1y agoTo me, the most implausible thing about ZFS-without-ECC doomsaying is the presumption that the failure mode of RAM is a persistently stuck bit. That's way less common than transient errors, and way more likely to be noticed, since it will destabilize any piece of software that uses that address range. And now that all modern high-density DRAM includes on-die ECC, transient data corruption on the link between DRAM and CPU seems overwhelmingly more likely than a stuck bit.