11 ms·
File Systems, Data Loss and ZFS
- ferrantim 12y agoThanks for the explanation of misdirected writes. I've heard the term before, but didn't know exactly what caused it. Reading this post was like watching one of those How Things are Made shows on the Discovery Channel. Very interesting to see how some things I take for granted actually work.
- ryao 12y agoMisdirected writes are not as well known as they should be. I am happy to increase awareness of them.
- oakwhiz 12y ago>ZFS is operating on a system without an IOMMU (Input Output Memory Management Unit) and a malfunctioning or malicious device modifies its memory. If a Linux system possessing an IOMMU was booted with iommu=pt as a kernel command line option, does the IOMMU still protect from this type of failure? This option puts the IOMMU into passthrough mode which is required to successfully use peripherals on some motherboards.
- ryao 12y agoNo. This mode was introduced specifically for virtualization so that the IOMMU will only restrict access to a guest machine's memory, such as when KVM is in use: http://lwn.net/Articles/329174/ http://lwn.net/Articles/329174/ The only case in which this would help is when ZFS is on the host and a device passed through to a guest malfunctions.
- IgorPartola 12y agoI found the Reordering Across Flushes section really interesting. So one rule of thumb is that you should not use hardware RAID with battery backup? Are there other types of hardware that would give you the same problems?
- mnw21cam 12y agoAs with everything, It Depends. Battery backup with a hardware RAID controller is fantastic for increasing write performance of a RAID array. You can literally call flush() on a file, and it returns almost instantly, so if you are doing this a lot it makes sense to either use one of these or a solid state drive. If the system is well-maintained, then the chances of it failing are "small". However, the whole ethos of ZFS is that it wants to manage individual discs itself. With regard to other hardware, there are certainly devices out there that will ignore flush requests, in the name of performance. These are usually consumer-grade devices. If you're paying extra for an enterprise-grade drive from a reputable manufacturer, you should be fine.
- ryao 12y agoI usually advise people to avoid hardware RAID controllers on the basis that they introduce unnecessary risk. If you want speed, it is best to use a device like the ZeusRAM as a SLOG device: http://www.hgst.com/solid-state-storage/enterprise-ssd/sas-ssd/zeusram-sas-ssd http://www.hgst.com/solid-state-storage/enterprise-ssd/sas-s... As for other devices, it is possible for bugs in software block devices to cause reordering across flushes. Finding out requires reviewing the code for any way that an IO before a flush can occur after it. This is a better situation than that with hardware RAID controllers, whose firmware is closed source and cannot be inspected.
- the8472 12y agoGenuine non-volatile RAM would avoid the issue of battery failures. About two years ago LSI announced a partnership to use MRAM in their RAID controllers but I haven't seen a product materialize out of that.
- mbreese 12y agoWell, with ZFS you want to avoid hardware RAID controllers completely. The protections from ZFS only work if the filesystem doesn't have anything in between it and the actual disks. Depending on your vendor, it can actually be difficult to get a card that lets you have JBOD access to a large disk array. The only exception that I can think of is encryption. You could wrap a disk with an encryption layer in software, but then you could still to make a separate virtual device for each disk.
- contingencies 12y agoTLDR: "its data integrity capabilities far exceed any other production filesystem available on Linux today"
- pedrocr 12y agoDoes anyone have a good up-to-date comparison with btrfs on this topic?
- xenophonf 12y agoCheck out the previous HN thread on ZoL: https://news.ycombinator.com/item?id=8303333 https://news.ycombinator.com/item?id=8303333
- pedrocr 12y agoThanks that was very helpful. Seems btrfs is still a bit behind, some of it by design. Valerie Aurora, who worked both on ZFS and btrfs seems to think the btrfs architecture is better in a few ways: https://lwn.net/Articles/342892/ https://lwn.net/Articles/342892/
- xenophonf 12y agoThanks for the link. That was a very interesting article. Btrfs sounds very interesting.
- lmm 12y agoI switched to FreeBSD a couple of years ago, partly for the sake of ZFS which is a first-class filesystem on that platform. FreeBSD was much more similar to Linux than I expected, and where there were differences, the FreeBSD way was usually simpler. My system has been stabler ever since, and I no longer fear to hit the "update" button.
- laumars 12y agoYeah I love FreeBSD. I wish more people gave it a chance before rushing to ZFS FUSE. Granted things are a little different now that ZoL is around and has proven to be stable, but it always struck me as a little odd that some would point blank refuse to even try FreeBSD yet welcome the lesser tested and poorer performing solution of running ZFS in FUSE. But each to their own I guess.
- existencebox 12y agoI was one of the people who rushed to do a ZFS setup on Ubuntu when those capabilities first started appearing ~5 or so years ago. There were some strange bugs that pushed me onto BSD, and the entire time since then since then I've been so stable it almost makes me nervous to try again despite the positive reception of modern ZOL (if it aint broke etc). Seriously, impressively stable.
- ferrantim 12y agoIf you haven't seen it already, this post by the same author about the State of ZFS on Linux might interest you: https://clusterhq.com/blog/state-zfs-on-linux/ https://clusterhq.com/blog/state-zfs-on-linux/
- existencebox 12y agoI've read it, but thanks for the breadcrumb. It definitely seeded thoughts in my head about giving it another try the next time I do a clean reformat on my fileserver.
- 12y ago
- lsllc 12y agoZFS on CoreOS anyone? [CoreOS does have btrfs support]
- phireph0x 12y agoDoes ZFS on Linux support ARM? I'd like to give it a spin in Arch Linux ARM.
- tkinom 12y agoI like to see that too! I worked on ARM Linux NAS system for an SOC company for a while. I think ZFS on ARM is good idea mainly because the low cost, low power CPU. Does anyone else see such need? If so, fill out this survey: https://docs.google.com/forms/d/1HRt_aYmuGkyQvBp9p2Tr3-9ygXR0YyTe2u7rtjTirOo/viewform?usp=send_form https://docs.google.com/forms/d/1HRt_aYmuGkyQvBp9p2Tr3-9ygXR... I am trying the lean startup method. :-) If there is < 20 people show interested in this concept, I won't spend more time on it.
- mbreese 12y agoHave you thought about FreeBSD for the NAS? It's a pretty common base for home-brew NAS systems. I'm not sure how the ARM support is though.
- ryao 12y agoIn theory, yes, but in practice, 32-bit support is a work in progress. You could try it, but you would want to make certain that you boot the kernel with vmalloc set to something larger than the amount of RAM that the system has, yet smaller than the 2GB of kernel address available on 32-bit (e.g. 1G). Otherwise, you could run into problems where kernel virtual memory allocations hang because of virtual address space exhaustion. This is due to a design decision in Linux to cripple kernel virtual memory, although it does not affect 64-bit systems because the kernel virtual address space is much larger than system memory at this time. This should be fixed in the next 6 months, but until then, you will need to be careful with it.
- deleted 12y ago[deleted]
- Someone 12y ago"In the case that we have two mirrored disks and accept the performance penalty of the controller reading both, the controller will be able to detect differences, but has no way to determine which copy is the correct copy." If you 'seed' the checksum algorithm for a block with the block number being written, a subsequent read of a different block that produces the same data will have a checksum failure. That would make it possible to choose which block has the right data. So, if you are willing to eat the performance, you can detect single misdirected writes.
- ryao 12y agoWhen I wrote that, I was talking about hardware RAID 1, which has no checksums.
- Someone 12y agoBut the disks have their own checksums, haven't they?
- ryao 12y agoThe low level formatting has ECC, which never leaves the drive. That said, there are two cases to consider for misdirected writes. One is that the write clobbers multiple sectors in which case you would get uncorrectable sectors. The second is that it perfectly replaces another sector. In that case, the ECC is a perfect match as the ECC is stored with the sector. Neither drive would report a problem, but the data would not match. This is what I described as being a problem and traditional RAID is incapable of dealing with it.
- Someone 12y agoAnd that is where I stated that drives can report a problem. If they 'seed' their ECC algorithm with the sector number (XOR-ing the result with it would be sufficient), they can (statistically) detect that, when they read sector #X, what they got wasn't what they ever wrote as sector #X. In fact, I guess they already do. If they didn't, there would be misdirected reads, too.