3 ms·
Why does ZFS "really need" ECC? I have a few friends running freenas on just old machines without ECC, no issues so far a few years in.
by rightos 9y ago
Why does ZFS "really need" ECC? I have a few friends running freenas on just old machines without ECC, no issues so far a few years in.
- slobotron 9y agoIf anything, ZFS checksumming lets you get by without ECC and still be confident in your data...
- kobeya 9y agoNo because the checksums are cached in ram and not verified against the disk when an error occurs.
- slobotron 9y agoThanks for catching that - my statement was definitely too strong in implying data invulnerability. I was hinting that you still get a bit of extra protection with ZFS over other filesystems, both on Non-ECC RAM. There is a chance with a stuck/corrupt bit that checksum will fail and you will get a read error. I interpreted OP as saying ECC is a must for ZFS, but I don't think you are more prone to corruption that any other FS? I'm basing my understanding on this: http://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-your-data/ http://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-yo...
- raattgift 9y agoYes-and-no. In the most general case the checksum of data block A lives in the parent block B. B's checksum in turn lives in its parent block. And so forth all the way to the root of the Merkle tree. These checksums are fixed exactly once, in the zio pipeline, as a dirtied buffer is being passed through it during a write operation. The checksum then stays on disk in that form in block B for as long as block A is reachable. In the read case, checksum validation is done at the vdev_{raidz,mirror,...}.c level and repaired there if possible e.g. from another element of the mirror; recovery is also possible in the case of copies=N (N > 1). Generally, the checksums of read data and metadata are not used once a buffer is successfully in the ARC. (However, compare arc_buf_{freeze,thaw}(), which one can enable, and which will happily cause panics if ARC buffers are corrupted whether through physical problems (e.g. in memory, in a bus, in the CPU) or software error in the kernel.) Block B may already be in cache for some other reason, in which case the in-memory checksum is used (this is the "yes" part of "yes-and-no"); otherwise, block B will have to be fetched before block A, because block B is where A's checksum lives. Block B may be evicted from ARC long before block A is evicted; any read() or similar call that needs data from A will not result in B being read back into memory. A buffer kept in ARC and used as the source of subsequently-dirtied data will have its checksum generated anew, just like the fresh write above. Obviously if the data it is dirtied with is bad, it will go to disk and its parents will have a checksum reflecting the bad data. However, the source of the badness need not be within the transient dirtied buffer, nor in a long-resident ARC buffer; userland memory can be corrupted too. If one has bad memory than bad data may end up on disk, checksummed in such a way that the badness will not be detected by the ZFS subsystem. But this is really no different from any other RAID-like data-validating system or a traditional one-disk filesystem which doesn't checksum. If bits are flipped in userland data that is ultimately subjected to a write(2) call (etc., think mmap()), no filesystem can reasonably expect to do anything other than write out the data handed to it via the system call. Like any other filesystem, ZFS has its own metadata for tracking allocations, object metadata, and in generating the branch-to-root Merkle tree. The wrong memory corruption (and this has happened in software during ZFS development) can result in data loss or even a pool too damaged to be imported (go to backups). ZFS is not especially more exposed to that than any other filesystem. Since the arrival of compressed ARC and crypto, only a small fraction of ARC buffers are kept uncompressed or unencrypted in memory at any time (this is controlled by the dbuf cache tunables). Bitflips in any encrypted-in-memory ARC buffer will be caught (noisily) if the buffer is subsequently read or modified, since a decryption will be done at that time and would fail under a single modified bit. Many possible corruptions in compressed-in-memory data will also be caught depending on the type of compression used when the buffer was first written out to disk. In neither case is the ZFS checksum involved in the catching of such corruptions. Finally, this still won't protect against corruptions in userland, corruptions which hit the various data structure that point to buffers in ARC, or corruptions in the text segments in the zfs subsystem or elsewhere in the kernel. However, ZFS is not realisically more fragile to such corruptions than any other filesystem.
- milcron 9y agoTraditional filesystems will let files just sit on the hard drive, untouched. ZFS is a lot more active... checksumming and caching mean that important information is spending time in RAM. It's not good if that information gets corrupted.
- rightos 9y agoDoesn't ZFS perform checksumming and error correction on that data to make sure that memory or disk issues don't get affected by corruption? I thought that was a main reason to use it.
- milcron 9y agoWell right. It protects against silent bitrot on the hard drives. But it does this by caching and checksumming in RAM, so now you need to defend against bitflips in RAM. It's not that ZFS without ECC is completely unsafe... it's just not as safe as it could be. ECC RAM becomes the next biggest concern once you've addressed on-disk bitrot.
- Dylan16807 9y agoChecksumming doesn't make data loss worse. ZFS doesn't cache in any special way, or more than your typical filesystem. Lack of ECC just lessens the benefits of ZFS, it doesn't exacerbate any problems.