6 ms·
Care to elaborate?
by ferrantim 12y ago
Care to elaborate?
- valarauca1 12y agoZpool is a great feature. On top of the dozen(s) of amazing features in ZFS. The problem is, and ZFS admits this: If you don't run ECC ram Zpool can "accidentally" your whole hard drive (bit of a joke there, it can corrupt your drive as it attempts to correct bit rot that never happened). This happens a lot more often then we really care to think about. (Ram corruption) For most day-to-day linux users who are just using standard consumer desktop/laptops they don't have this protection. And ZFS can do more harm then good. References: https://pthree.org/2013/12/10/zfs-administration-appendix-c-why-you-should-use-ecc-ram/ https://pthree.org/2013/12/10/zfs-administration-appendix-c-...
- ryao 12y agoThis is a myth started by someone who did not understand how filesystems work. Not having ECC memory is no more of a handicap for ZFS than it is for any other filesystem.
- valarauca1 12y agoCare to elaborate? ZFS warned me of that while I reading its installation guide. Which is why I never installed it.
- sp332 12y agoIf you're doing RAID, you should have ECC. All hardware RAID controllers have ECC cache, and if you're doing software RAID, you should have ECC system RAM. But this is just common sense, I guess, since it's useless to write data to redundant storage if it's already been bit-flipped in RAM.
- ryao 12y agoI thought that I had replied to this, but I seem to have replied to its parent by mistake. Here is a link to the response: https://news.ycombinator.com/item?id=8438416 https://news.ycombinator.com/item?id=8438416 Would you provide a link to that guide? If there is a guide out there that says that you should use something other than ZFS when a system lacks ECC, I would like to know so that I can try to get it corrected. Not using ZFS because it cannot provide full protection without ECC is like not getting a flu shot because you can still get sick anyway. It is true, but opting to be even less safe in the name of safety is counterproductive.
- sp332 12y agoThe trouble is, once you have a corrupted ZFS, there's no good way to recover it. That's why it's vital to make sure nothing bad happens in the first place.
- ryao 12y agoThe same can be said for other filesystems. The issues that fsck has been abused to automatically fix on them simply do not happen on ZFS. Failures so severe that they kill ZFS have should have analogous failure states on those filesystems too. People do not hear about such failures because they cannot be distinguished from the more typical issues that affect such filesystems.
- moe 12y agoThis is a myth started by someone who did not understand how filesystems work. You are wrong. Don't spread potentially dangerous maladvice on the internet when you have no idea what you're talking about. The official ZFS documentation[1] tells you to use ECC Ram and why. The first google hit for "zfs ecc ram"[2] further elaborates on the risks of using ZFS without ECC memory. [1] https://pthree.org/2013/12/10/zfs-administration-appendix-c-why-you-should-use-ecc-ram/ https://pthree.org/2013/12/10/zfs-administration-appendix-c-... [2] http://louwrentius.com/please-use-zfs-with-ecc-memory.html http://louwrentius.com/please-use-zfs-with-ecc-memory.html
- skorgu 12y agoRAID is not a backup, zfs especially so. Zfs (and RAID) is designed to protect your data from a small, enumerable and specific set of failures, namely the right-there-in-the-acronymn "inexpensive disks". If your house burns down ZFS will not save your data. If too many disks fail zfs will not save your data. If your disks' unrecoverable read error rate is too high and your array is rebuilding zfs will not save your data. If your computer is fundamentally broken (which is what a machine with a persistent memory error is) zfs will not save your data. Take backups. If you had a 'stuck bit' anywhere in your memory space you'd be way deep into nasal demon unspecified behavior the first time you tried to dereference a pointer that crossed that bit. When your hardware is that broken you can't count on the OS to not stab your dog much less ZFS to do anything sane. Note that this is just as true for all storage systems. Regular filesystems might corrupt themselves or do other insane things. Hardware RAID or kernel software raid will happily propogate the error. How many bits separate the kernel record for "this disk is totally cool" from "this is a new disk and should be zeroed"?
- ryao 12y agoWhile it is possible to have an unimportable pool, it is also possible to have ext4 and XFS filesystems that can neither be mounted nor repaired with fsck. When dealing with undefined behavior caused by bit flips, virtually any failure is possible, especially when you consider bit flips to kernel data structures. No software can save you from bit flips and if they are your concern (as they should be), then you should refuse to use computers that lack ECC RAM. As for the articles that you link, they are correct to say that you want to use ECC RAM. However, there is nothing specific to ZFS that makes it require ECC any more than any other filesystem. It should also be noted that I wrote the official documentation on this subject. It can be found at the Open ZFS wiki, rather than the pages you linked: http://open-zfs.org/wiki/Hardware#ECC_Memory http://open-zfs.org/wiki/Hardware#ECC_Memory
- ryao 12y agoA person at the FreeNAS forums wrote a forum post explaining his mistaken believed that ZFS is somehow more prone to catastrophic data loss when bit flips occurred than other filesystems: https://forums.freenas.org/index.php?threads/ecc-vs-non-ecc-ram-and-zfs.15449/ https://forums.freenas.org/index.php?threads/ecc-vs-non-ecc-... The reality is that the worst case consequences for a bit flip is the complete loss of all data, regardless of the filesystem used. While the automated repair routined bundled in fsck utilities are often able to fix problems on other file systems, the problems that they do fix simply do not occur on ZFS. The kinds of problems that kill ZFS are not among those that an automated repair tool can fix. e.g. overwrite all ext4 superblocks with random data and then see if fsck.ext4 can fix it. That said, I am usually able to resuscitate a pool that another person would have considered to have been killed by a bit flip. I not only find such failure modes to be incredibly rare, but I find those that I cannot fix to be a rarity among the cases where things did in fact go wrong.
- atoponce 12y agoIndeed. I can reiterate this point. I have helped a few administrators, who initially were compiling ZFS on Linux from source, then decided to switch to the Launchpad PPA repository, and "upgrade" their userspace tools. Next thing I know, they are asking me if there is anything they can do to recover their pool. After a bit of research, and a little cleaning up, I am able to get their pool back online, with zero data loss. I have had a VERY hard time finding a situation where there was corrupted data in ZFS, and where there are ZFS pools that absolutely will not import back into full operation. In other words, it is pretty difficult to "accidentally" corrupt a ZFS pool to the point of not being able to recover your data. ECC RAM greatly minimizes the potential for a corrupted file in ZFS, but ECC RAM also greatly minimizes the potential for a corrupted file in ext4 or XFS. So long story short, I have experienced the same thing as ryao, and come to the same conclusions.
- po 12y agoAre you referencing this thread? https://groups.google.com/forum/#!topic/zfs-macos/qguq6LCf1QQ https://groups.google.com/forum/#!topic/zfs-macos/qguq6LCf1Q... I've heard this but I haven't actually confirmed anywhere official that ZFS will try to scrub your data due to a parity flip in ram values. Where in the manuals do they talk about it? (I'm actually curious about if it's just a rumor or if it's acknowledged)
- valarauca1 12y agoReferencing this https://pthree.org/2013/12/10/zfs-administration-appendix-c-why-you-should-use-ecc-ram/ https://pthree.org/2013/12/10/zfs-administration-appendix-c-... http://louwrentius.com/please-use-zfs-with-ecc-memory.html http://louwrentius.com/please-use-zfs-with-ecc-memory.html
- po 12y agoEven after reading that, I'm still not convinced it's any worse than something like btrfs+non-ECC memory. The most convincing part is the 'there are less tools' to recover from corruption angle.
- Thaxll 12y agoThe article is missing many things that make a FS suitable for production, it's listing every features but lack the reliability / cons / waknesses ect... ZFS is nowhere near production ready for Linux. I know XFS and ext4 under heavy load with different scenarios, could you tell the same for ZoL?
- atoponce 12y agoYes, I can. I am a system administrator who also knows a great deal about storage, and manages a good chuck of the storage servers at my employment. We use ZFS for our backup servers, both onsite, and offsite, and we use them for a couple generic storage servers as well, one of which is constantly under heavy stress, all the time. That server is http://mirrors.xmission.com http://mirrors.xmission.com. I am personally using it on my workstation for a /home mount, and I have it on a highly available KVM 2-node cluster, replicated with GlusterFS using InfiniBand. While I have had networking issues with GlusterFS, which have since been ironed out, I have not had any stability, reliability, or data corruption issues with ZFS. At all. I've been running this cluster for 3 years straight, and it's the cluster that is housing my ZFS documentation at https://pthree.org/category/zfs https://pthree.org/category/zfs ZFS on Linux is absolutely "production ready". While it's true there are some ARC issues floating about, trim support is missing for SSDs, and some other things, remember that ZFS has been stable for a decade. It's just been brought into Linux kernel space with the help of the Solaris Porting Layer (spl) module. Claiming that ZFS on Linux is not stable, and not production ready is nothing more than FUD.
- mh- 12y agois lack of trim support not a substantial issue? (for those using SSDs)
- atoponce 12y agoIt can cause performance degradation when all the banks in the SSD have been written to, and a new bank needs to be re-written. TRIM allows cleaning the bank when it is no longer needed, and hopefully, before it needs to be written again. As such, the new rewrite will be native speed, and not bottle-necked