15 ms·
ZFS for Dummies
- dontupvoteme 3y agoAs nice as the technology is as long as there's the potential of a Damoclean license issue I'll always feel hesitant around ZFS. (Hey looks like it's a sore spot!)
- judge2020 3y agoEven in such an event, I don’t see how it would affect end users in an extreme way. OpenZFS doesn’t call home to make sure it’s still legal to use, so worst that could happen is that they are forced to stop distribution and you need to stay on old software until you can move to another storage system.
- NavinF 3y agoNot to mention Ubuntu included ZFS by default since 2016. I still remember all the doomsday predictions people made back then. People on HN and proggit kinda suck at legal
- arjvik 3y agoPlaying devil's advocate here, but what if Oracle decides to sue you, a company with only a few thousands in revenue, for their multi-million dollar licensing fee? I agree this is unlikely, but so is someone being born with as much litigiousness as Larry Ellison.
- TheDong 3y agoThis is built on a misunderstanding of what the ZFS licensing issue is. ZFS is licensed under the CDDL [0], which is a copyleft license that grants you full rights to use the ZFS source code for basically anything, providing that any derivative works incorporating its source code are also are under the CDDL. There is no "licensing fee" that you can pay Oracle related to this. Linux is licensed under the GPL, which just like the CDDL, is a copyleft license which requires derivative works to be distributed under the GPL. The conflict here is that both require the entire derived work to be distributed under CDDL/GPL, but obviously only one license can be picked.... even though realistically, GPL and CDDL are quite similar licenses. The additional fact that is useful to know is that zfs.ko, the zfs kernel module, is distributed under the CDDL by ubuntu, and any other distro distributing it that I know of. So, with that knowledge, who can sue who and why? 1. The people that can be sued are people distributing compiled zfs binaries (the 'zfs.ko' kernel binary). In practice, this means canonical, and anyone else who distributes disk images using ZFS (such as snapshotted EC2 AMIs) 2. The people who can sue are the holders of the Linux GPL license, and they can do so by claiming zfs.ko violates their copyright by being a derived work of the linux kernel, but being distributed under the CDDL rather than the GPL. So, first of all, the vast majority of companies using ZFS aren't also distributing binary distributions containing zfs.ko and the linux kernel together, therefore this whole thing is moot. Realistically, most companies using ZFS only have the risk of _canonical_ being sued, and that breaking their filesystem, not them running into legal issues themselves. And second, it would require a holder of Linux copyright to file suit, and realistically, no court could find significant damages here. Canonical also has lawyers that claim 2 wouldn't legally follow, that 'zfs.ko' is not in fact a derived work of the linux kernel. There's differing opinions there, but pretty much everyone agrees that the people who could sue if this interpretation is wrong is the holders of Linux copyright, not oracle. [0]: https://en.wikipedia.org/wiki/Common_Development_and_Distribution_License https://en.wikipedia.org/wiki/Common_Development_and_Distrib...
- arjvik 3y agoI see, thanks for educating me on this with such an in-depth reply! I simply saw "licensing issue" and "Oracle" together and assumed the standard Oracle licensing worry applied.
- jclulow 3y agoNote that unlike the GPL, the CDDL is file-based. It doesn't require your entire body of software to be CDDL, it just requires that you make modifications to CDDL-licenced files available under the CDDL.
- jakobson14 3y agoYou have one crucial fact wrong: The CDDL is actually a weak per-file copyleft. The CDDL is totally ok with the rest of the linux kernel staying GPLv2. The GPL is allergic to any non-GPL (eg. ZFS) code being distributed with the kernel source.
- TheDong 3y agoSure, but I don't think that's crucial here. Does it change any of the other information? I admit I shouldn't have written "both require the entire derived work to be distributed under CDDL/GPL", and instead should have written "both may require, depending on a court's interpretation, the zfs kernel module to be distributed under CDDL or GPL respectively". I think the rest of the comment stands though, right?
- jakobson14 3y ago> "both may require, depending on a court's interpretation, the zfs kernel module to be distributed under CDDL or GPL respectively". That's wrong, and misses fundamental information about the CDDL. First, the CDDL gives absolutely no fucks how you licence a resulting binary. It only cares about the specific covered source files, and doesn't care what you bundle them with. So nobody would ever argue that the module should be licenced as CDDL. On the flip side, you'd have a really hard time arguing that the module should be licensed as GPL.* Sure, there's maybe a GPL wrapper, but lots of modules are non-GPL even when they might feature such a wrapper. ie. nvidia. You also can't argue ZFS is a derived work of the kernel, which is really the crux of the issue. ZFS was first written on solaris, developed for several years on freeBSD, and remains neutral across 3-5 different supported OSs. There's nothing to indicate the GPL should have more claim over ZFS than the wrapper. The one and only place you'll ever run into trouble is if you start trying to distribute ZFS *source code* in the kernel source tree, as a fully integrated part of the kernel. (EDIT: or compiled together in a single kernel binary, from said combined source) The CDDL doesn't care about the mixing but the GPL does. Literally everyone agrees this is a bad idea and nobody does it. *Let's ignore for a moment for the moment that you *couldn't* since the GPL requires complete corresponding source code and you can't relicence the CDDL source code files.
- kaba0 3y agoOn what grounds? ZFS has an open-source license. The only semi-open question related to the topic is packaging it together with another, different open-source license that is viral and mandates more things (GPL). That part is not being “violated” by an enduser, but the distro maintainer.
- jakobson14 3y agoOracle's hands are tied in so far as they can't revoke the CDDL licence that Sun provided. They could sue you as a kernel contributor, but only if you do something to violate the GPL. The only thing that would clearly violate the GPL, in the era of proprietary nvidia blobs, is if you tried to ship the ZFS source code in-tree. Everybody agrees that's a bad idea, nobody does it. ZFS instead is shipped as a separate kernel module.
- wkat4242 3y agoIt's not really an issue with ZFS itself. It's a great filesystem. GPL isn't the only way to do open source. It's the way Linux has chosen and that has pros and cons. Not integrating so well with other open source licenses is one of the cons.
- tylercrompton 3y ago> Not integrating so well with other open source licenses is one of the cons. That depends on what you want. If you want a license that will play well with closed source software, then yeah, it's a downside. But the GPL family comes from the perspective of a developer who wants to retain their rights while respecting others' desire for the same. If you care about your rights, then this is an upside.
- kaba0 3y agoCDDL (ZFS’s license) is also like that, the problem is exactly the two “strong” copyleft licenses.
- jakobson14 3y agohttps://www.gnu.org/licenses/license-list.en.html#CDDL https://www.gnu.org/licenses/license-list.en.html#CDDL >"It has a weak per-file copyleft (like version 1 of the Mozilla Public License)" The CDDL isn't a strong copyleft licence. It doesn't give a shit what larger project you include the code into, and it doesn't give a shit how you licence the resulting binary. It's considerably more permissive than the GPL. The conflict is entirely on the GPL side.
- ggm 3y agoThis is only going to get a shedload of "you do you" responses. BSD licencing is quite content to use ZFS, its pretty much ubiquitous now in widescale deployment and the source would be impossible to re-lock. Worst case is fork. I very much regret the fragmentation of FS design, it has many mothers. "there can only be one" was never going to work, but we seem to have perhaps 4-5 more than we really need. ZFS manages to wrap up a number of behaviours cohesively with good version-dependent signalling so it should always be possible to know you're risking a non-reversible change to your flags. And, it keeps improving. But, counter "it keeps improving" so do all the other current, maintained, developed FS and if somebody tells me they prefer to use Hammer, or one of the Linux FS models with a discrete volume-management and encryption layer, I don't think thats necessarily wrong. Mainly I regret Apple walking away. That was about Oracle behaviour. It wasn't helpful. A lot of Apple's FS design ideas persist. I never got resource/data forks, it only ever appeared on my radar as .files in the UNIX AUFS backend model of them. Obviously inside Apples code, it was dealt with somehow. It felt like the wrong lessons about meta data had been learned. Maybe an Ex-VMS person went to Apple? Also Apple has a rather "yea maybe or no, dunno" view about case-independent or case-dependent naming. Time machine is good. Feels like it should fit ZFS well. Oh well.
- chaxor 3y agoI wish zfs had a 'tag' type of feature like Apple style fs. I have thought about making an object store to get some of the metadata on the fs level, rather than having to make some cobbles together solution of an SQLite DB of metadata with the files or something, but many of the object store solutions seem to be for fairly large files, rather than a few thousand pictures of family and such. Unsure what a good solution using zfs would be.
- adrian_b 3y agoI am not sure what kind of "tag" you mean, but in any modern file system which supports extended file attributes you can attach arbitrary metadata to any file. You just have to define whatever kind of tags you want and choose names for the xattrs in which to store them and then you should use them in your scripts or applications. I am using all the time xattrs for things like file checksums and content classification, on XFS and on FreeBSD UFS, but almost all modern file systems support xattrs, except Linux tmpfs (which has only a partial support that is useless for those who use custom xattrs, so a file with xattrs cannot be copied to /tmp if tmpfs is used).
- rodgerd 3y agoDo you refuse to use NVidia cards for the same reason?
- dontupvoteme 3y agoNo choice on that front.
- NavinF 3y agoNew drivers are MIT/GPL2 dual-licensed: https://github.com/NVIDIA/open-gpu-kernel-modules https://github.com/NVIDIA/open-gpu-kernel-modules
- mrweasel 3y agoDoesn't that still require the closed source user space driver to function? I'd love to be wrong on that one.
- vetinari 3y agoIn the short-term yes, in the long term the user space driver would also get open (or replaced with nouveau, since it will be on equal footing with the closed one, when communicating with kernel; just like the open vs closed amd drivers). However, what will be left is a blob, that runs on the dedicated core on the GPU. The stuff that Nvidia doesn't want to open is moved there, and that's why the open kernel driver works only on Turing and newer.
- lmm 3y agoYou might want to consider using FreeBSD then, where there is no licensing issue. I've found it has all the stuff I liked about Linux, and less of the stuff I dislike.
- seized 3y agoThere is also Illumos/OpenIndiana/OmniOS, basically continuations of OpenSolaris. I have used OpenIndiana for many years and it's been dead reliable and stable. There's quite a few "quality of life" differences like boot environments (boot into a pre upgraded OS state, even years old), built in SMB server with NFS v4 style ACLs, dtrace, built in snapshot scheduling and management, and Napp-It is an available web UI for management a la FreeNAS/TrueNAS. It has a few differences, service management is quite different from other things, but overall very underrated as an OS I think.
- unethical_ban 3y agoI haven't looked at those distributions in years. I thought they were niche, container host operating systems and not general purpose.. I guess it need to do my research.
- kaba0 3y agoThe license issue in the very worst case (but even that is quite questionable and I would say that Linus is a bit paranoid here) could only be a problem for a distro that ships it by default, you as the end user can freely use both linux and ZFS, as well as their combination.
- tomxor 3y agoI was concerned about this aspect initially, but ZoL has been going for many years now. At this point even Debian (the most purist in terms of GPL considerations) includes ZoL in it's official repository. It doesn't distribute binaries, apt just builds the kernel module automatically. Frankly, if even Debian can use it, it's a non-issue.
- unethical_ban 3y agoZFS on Linux, and even Ubuntu distributing the binary form in their base OS, has been going on for years. I think you would be a huge event and very unlikely to occur if Oracle tried to pull some licensing bullshit. As someone ignorant to file system development, I would almost expect something more likely to be BTRFS getting sued for copying a feature of ZFS or something like that.
- jakobson14 3y agoThere's really no case to be made from the CDDL side. It's a weak per-file copyleft. If anyone (Oracle or Linus Torvalds) launches a ZFS-related lawsuit, it'll be as an author of GPLv2 kernel code. For the time being the solution has been to ship ZFS separate from the kernel, as any module with a non-GPL licence typically does. The biggest hurdle to ZFS isn't legal, but technical. The various teams working on ZFS has so far been able to keep up with kernel churn and symbols (eg FPU ones) being made GPL-only. That said, the kernel devs have made it clear that they don't care if you're open source or proprietary, they will make changes and mark new symbols GPL-only to fuck with you regardless.
- totetsu 3y agoNice. My gotchas form using zfs on my personal laptop with Ubuntu. - if you want to copy files for example and connect your drive to another system and mount your zpool there, it sets some pool membership value on the file system and when you put it back in your system it won’t boot unless you set it back. Which involved chroot - the default settings I had made snapshot every time I apt installed something, because that snap shot included my home drive when I deleted big files thereafter I didn’t get any free space back until i figued out what was going on and arbitrarily deleted some old snapshots - you can’t just make a swap file and use it,
- Helmut10001 3y ago> - if you want to copy files for example and connect your drive to another system and mount your zpool there, it sets some pool membership value on the file system and when you put it back in your system it won’t boot unless you set it back. Which involved chroot Isn't this what `zpool export` is for?
- dizhn 3y agoOpensuse Tumbleweed comes with snapper which works with btrfs in a similar fashion and /home is not included in the snapshot by default. For your use case too you should exclude /home from your apt triggered snapshots and set a separate one for it. I had scheduled snapshots for my /home at one point but since it's a very actively used directory (downloading isos, games then deleting them) I had similar problems to yours. I guess we could both also have a dedicated separate directory for those short lived huge files, which don't need a snapshot anyway.
- Dylan16807 3y ago> I had scheduled snapshots for my /home at one point but since it's a very actively used directory (downloading isos, games then deleting them) I had similar problems to yours. What kind of schedule was it? I feel like the low-impact alternative to no snapshots at all is daily snapshots for half a week to a week, and maybe some n-hourly snapshots that last a day or two. Which I would not expect to use up very much space.
- asicsp 3y agoSee also: https://jro.io/truenas/openzfs/ https://jro.io/truenas/openzfs/ (discussed here: https://news.ycombinator.com/item?id=34318734 https://news.ycombinator.com/item?id=34318734)
- guerby 3y agoI started to use ZFS (on Linux) a few years ago and it went smoothly. My only surprise was volblocksize default which is pretty bad for most RAIDZ configuration: you need to increase it to avoid loosing 50% of raw disk space... Articles touching this topic : https://jro.io/nas/#overhead https://jro.io/nas/#overhead https://openzfs.github.io/openzfs-docs/Basic%20Concepts/RAIDZ.html https://openzfs.github.io/openzfs-docs/Basic%20Concepts/RAID... https://www.delphix.com/blog/zfs-raidz-stripe-width-or-how-i-learned-stop-worrying-and-love-raidz https://www.delphix.com/blog/zfs-raidz-stripe-width-or-how-i... And you end up on one of the ZFS "spreadsheet" out there: ZFS overhead calc.xlsx https://docs.google.com/spreadsheets/d/1tf4qx1aMJp8Lo_R6gpT689wTjHv6CGVElrPqTA0w_ZY https://docs.google.com/spreadsheets/d/1tf4qx1aMJp8Lo_R6gpT6... RAID-Z parity cost https://docs.google.com/spreadsheets/d/1pdu_X2tR4ztF6_HLtJ-Dc4ZcwUdt6fkCjpnXxAEFlyA https://docs.google.com/spreadsheets/d/1pdu_X2tR4ztF6_HLtJ-D...
- rhinoceraptor 3y agoIn my opinion, the 50% efficiency of mirror vdevs is a fair price to pay for the simplicity and greatly improved performance. You can grow RAIDZ pools now, but it's still a lot more complicated and doesn't perform as well.
- idatum 3y agoI need 3 stores to feel I'm keeping safe years of digital family photos. 1) I have a live (local) FreeBSD ZFS server running for backups and snapshots; 2 pairs of mirrored physical drives 2) I have a USB device that takes 2 mirrored drives to recv ZFS snapshots from #1; I store that vdev backup in a safe place 3) I backup entire datasets to cloud storage from off-prem using rclone. It's #3 where I need to do some more research/work. I need to spend some time sending snapshots/diffs to cloud blob storage and make sure I can restore. Yes, I know there is rsync.net. Any experiences to share?
- Maakuth 3y agoI have a similar setup, though with Linux. From my experience I can recommend taking a look at restic (https://restic.readthedocs.io/ https://restic.readthedocs.io/). It does encrypted and deduplicated snapshots to local and remote repositories. There's a good selection of remote target options available, but you can also use it with rclone to use any weird remote. Just remember to keep a backup of your encryption key somewhere besides the machine you back up ;-)
- deleted 3y ago[deleted]
- havnagiggle 3y agoMy setup is similar. +1 that Restic is great. My cloud backup was sending blobs to Google workspace, but they have clamped on storage. I will be replacing that with another box at my parents that will be tucked out of the way. I will just have that wireguard tunnel to my home network and send snapshots to it. At some point I'll turn down the workspace solution and probably also unsubscribe.
- zyberzero 3y agoI bought a cheap HP Microserver with four 4 TB spinning disks that I placed at a relatives house ~1000 km from where I live. I do nightly replication to the off-site location, with an account on the receiving end that only has enough permissions to create snapshots and receive data, so even if that ssh key somehow got out in the wild things could not be deleted from the remote store. I hope :) Clarification: Remote end also uses ZFS, so I can use cheap replication with encryption
- tomxor 3y agoI recently rebuilt a load of infrastructure (mainly LAMP servers) and decided to back them all with ZFS on Linux for the benefit of efficient backup replication and encryption. I've been using ZFS in combination with rsync for backups for a long time, so I was fairly comfortable with it... and it all worked out, but it was a way bigger time sink than I expected - because I wanted to do it right - and there is a lot of misleading advice on the web, particularly when it comes to running databases and replication. For databases (you really should at minimum do basic tuning like block size alignment), by far the best resource I found for mariadb/innoDB is from the lets encrypt people [0]. They give reasons for everything and cite multiple sources, which is gold. If you search around the web elsewhere you will find endless contradicting advice, anecdotes and myths that are accompanied with incomplete and baseless theories. Ultimately you should also test this stuff and understand everything you tune (it's ok to decide to not tune something). For replication, I can only recommend the man pages... yeah, really! ZFS gives you solid replication tools, but they are too agnostic, they are like git pluming, they don't assume you're going to be doing it over SSH (even though that's almost always how it's being used)... so you have to plug it together yourself, and this feels scary at first, especially because you probably want it to be automated, which means considering edge cases... which is why everyone runs to something like syncoid, but there's something horrible I discovered with replication scripts like syncoid, which is that they don't use ZFS's send --replication mode! They try to reimplement it in perl, for "greater flexibility", but incompletely. This is maddening when you are trying to test this stuff for the first time and find that all of the encryption roots break when you do a fresh restore, and not all dataset properties are automatically synced. ZFS takes care of all of this if you simply use the build in recursive "replicate" option. It's not that hard to script manually once you commit to it, just keep it simple, don't add a bunch of unnecessary crap into the pipeline like syncoid does, (they actually slow it down if you test), just use pv to monitor progress and it will fly. I might publish my replication scripts at some point because I feel like there are no good functional reference scripts for this stuff that deal with the basics without going nuts and reinventing replication badly like so many others. [0] https://github.com/letsencrypt/openzfs-nvme-databases https://github.com/letsencrypt/openzfs-nvme-databases
- yjftsjthsd-h 3y ago
- philsnow 3y agoOne of the diagrams under the bit about snapshotting has a typo reading "snapthot" and I immediately thought it was talking about instagram. (I realize now after writing it that maybe snapchat should have occurred to me first, but I have never used it)
- qwertox 3y agoFreeBSD's Handbook on ZFS [0] and Aaron Toponce's articles [1] were what helped me the most when getting started with ZFS [0] https://docs.freebsd.org/en/books/handbook/zfs/ https://docs.freebsd.org/en/books/handbook/zfs/ [1] https://pthree.org/2012/04/17/install-zfs-on-debian-gnulinux/ https://pthree.org/2012/04/17/install-zfs-on-debian-gnulinux...
- CTDOCodebases 3y agoI love FreeBSD's docs. I had an old HP Microserver with 1GB of ECC RAM lying around so I installed FreeBSD on it. I had 5 old 500GB hard drives lying around too so I set them up in a 5x mirror with help from the FreeBSD Handbook. First time using FreeBSD and it was a breeze.
- istjohn 3y agoI'm getting started with ZFS just now. The learning curve is steeper than I expected. I would love to have a dumbed down wrapper that made the common case dead-simple. For example: - Use sane defaults for pool creation. ashift=12, lz4 compression, xattr=sa, acltype=posixacl, and atime=off. Don't even ask me. - Make encryption just on or off instead of offering five or six options - Generate the encryption key for me, set up the systemd service to decrypt the pool at start up, and prompt me to back up the key somewhere - `zfs list` should show if a dataset is mounted or not, if it is encrypted or not, and if the encryption key is loaded or not - No recursive datasets and use {pool}:{dataset} instead of {pool}/{dataset} to maintain a clear distinction between pools and datasets. - Don't make me name pools or snapshots. Assign pools the name {hostname}-[A-Z]. Name snapshots {pool name}_{datetime created} and give them numerical shortcuts so I never have to type that all out - Don't make me type disk IDs when creating pools. Store metadata on the disk so ZFS doesn't get confused if I set up a pool with `/dev/sda` and `/dev/sdb` references and then shuffle around the drives - Always use `pv` to show progress - Automatically set up weekly scrubs - Automatically set up hourly/daily/weekly/monthly snapshots and snapshot pruning - If I send to a disk without a pool, ask for confirmation and then create a new single disk pool for me with the same settings as on the sending pool - collapse `zpool` and `zfs` into a single command - Automatically use `--raw` when sending encrypted datasets, default to `--replicate` when sending, and use `-I` whenever possible when sending - Provide an obvious way to mount and navigate a snapshot dataset instead of hiding the snapshot filesystem in a hidden directory
- deleted 3y ago[deleted]
- paulmd 3y agothis is honestly hard because many of the decisions that matter are not things you type into zfs at all (except incidentally). how many disks per vdev? how much memory? etc a lot of the things you've outlined are not universal at all, just situational
- istjohn 3y agoYes, there is a lot of essential complexity that is unavoidable, but there are a lot of people like me who just want a better desktop file system, and we don't need to know about SLOGs and L2ARCs and the half dozen compression algorithms, etc. It's situational, but it's a common enough situation that a targeted solution would be valuable.
- customizable 3y agoWe have been running a large multi-TB PostgreSQL database on ZFS for years now. ZFS makes it super easy to do backups, create test environments from past snapshots, and saves a lot of disk space thanks to built-in compression. In case anyone is interested, you can read our experience at https://lackofimagination.org/2022/04/our-experience-with-postgresql-on-zfs/ https://lackofimagination.org/2022/04/our-experience-with-po...
- photon_lines 3y agoNice - thanks for the info! I had no idea about the Toy Story 2 fiasco as well so this was a great read :)
- customizable 3y agoThanks, glad you liked it.
- PappGaborSandor 3y ago[dead]
- vermaden 3y agoOther useful things about ZFS: - get to know the difference between zpool-attach(8) and zpool-replace(8). - this one will tell you where your space is used: # zfs list -t all -o space NAME AVAIL USED USEDSNAP USEDDS USEDREFRESERV USEDCHILD (...) - ZFS Boot Environments is the best feature to protect your OS before major changes/upgrades --- this may be useful for a start: https://is.gd/BECTL https://is.gd/BECTL - this command will tell you all history about ZFS pool config and its changes: # zpool history poolname History for 'poolname': 2023-06-20.14:03:08 zpool create poolname ada0p1 2023-06-20.14:03:08 zpool set autotrim=on poolname 2023-06-20.14:03:08 zfs set atime=off poolname 2023-06-20.14:03:08 zfs set compression=zstd poolname 2023-06-20.14:03:08 zfs set recordsize=1m poolname (...) - the guide misses one important info: --- you can create 3-way mirror - requires 3 disks and 2 may fail - still no data lost --- you can create 4-way mirror - requires 4 disks and 3 may fail - still no data lost --- you can create N-way mirror - requires N disks and N-1 may fail - still no data lost (useful when data is most important and you do not have that many slots/disks)
- 65a 3y agoN-way mirrors also have the property that ZFS can shard reads across them, which mattered a lot on spinning rust, since iops can be limited.
- tweetle_beetle 3y agoMight not remember the details correctly but when I was younger and stupider I read a lot about how great one of the open source NAS OSs (FreeNAS?) and ZFS were from fervent fans. I bought a very low spec second hand HP micro server on eBay and jumped straight in without really knowing what I was doing. I asked a few questions on the community forum but the vast majority of answers were "Have you read the documentation?!" "Do you have enough RAM?!". The documentation in question was a PowerPoint presentation with difficult to read styling, somewhat evangelical language, lots of assumptions about knowledge and it was not regularly updated. It was vague on how much RAM was required, mainly just focused on having as much as possible. Needless to say I ignored all the red flags about the technology, the hype and my own knowledge and lost a load of data. Lots of lessons learnt.
- andruby 3y agoCan you roughly remember how long ago that was? ZFS has been around since the earlier 2000's, with FreeNAS starting in 2005 iirc. The filesystem has gotten a lot more stable, and imo the documentation clearer. That said, it's "more powerful and more advanced" than traditional journaling filesystems like ext3, and thus comes with more ways to shoot yourself in the foot.
- znpy 3y agois there an equivalent "btrfs for dummies" ?
- mastax 3y agoI've run into a ZFS problem I don't understand. I have a zpool where zpool status prints out a list of detected errors, never in files or `<metadata>` but in snapshots (and hex numbers that I assume are deleted snapshots). If I delete the listed errored snapshots and run zpool scrub twice the errors disappear and the scrub finds no errors. Zpool status never listed any errors for any of the devices. So there aren't any errors in files. There aren't any errors in devices. There aren't any errors detected in scrub(?). And yet at runtime I get a dozen new "errors" showing up in zpool status per day. How?
- Modified3019 3y agoDamn good question. I don’t have time to search for duplicates myself right now, but you can look through/ask the mailing list: https://zfsonlinux.topicbox.com/groups/zfs-discuss https://zfsonlinux.topicbox.com/groups/zfs-discuss (looks weird, but this is a legit web front end for the mailing list) and the github issues: https://github.com/openzfs/zfs/issues https://github.com/openzfs/zfs/issues
- unethical_ban 3y agoI've been running into the same issue, where occasionally files seem to get corrupted on the snapshot but also in the live version of the file. I cannot move it or modify it. I can only delete it. There's no indication as to why these files are getting corrupted. Thankfully there they are all large Linux ISOs, so it hasn't been critical to my life.
- footlose_3815 3y ago"Also read up on the zpool add command." Haha, The only part of maintenance that I need to look up every time I do it is replacing a faulty hard drive. Even this guide skips that.
- unethical_ban 3y agoSome additional points for posterity, in case it isn't driven home here: - All redundancy in ZFS is built in the vdev layer. Zpools are created with one or more vdevs, and no matter what, if you lose any single vdev in a zpool, the zpool is permanently destroyed. - Historically RAIDZs (parity RAIDs) cannot be expanded by adding disks. The only way to grow a RAIDZ is to replace each disk in the array one at a time with a larger disk (and hope no disks fail during the rebuild). So in my very amateur opinion, I would only consider doing a RAIDZ if it is something like a RAIDZ2 or 3 with a large number of disks. For n<=6 and if the budget can stand it, I would do several mirrored vdevs. (Again as an amateur I am less familiar with RW performance metrics of various RAIDs so do more research for prod).
- sgarland 3y agoPool of mirrors is usually the safer way, yes. If and only if you a. Have full, on-site backups b. Are fairly sure of your abilities and monitoring then I can suggest RAIDZ1. I have a pool of 3x3 drives, which ships its snapshots a few U down in my rack to the backup target that wakes up daily, and has a pool of 3x4 drives, also in RAIDZ1. In the event that I suffer a drive failure in my NAS, my plan of action would be to immediately start up the backup, ingest snapshots, and then replace the drive. That should minimize the chance of a 2nd drive failure during resilvering destroying my data. Truly important data, of course, has off-site as well.
- deleted 3y ago[deleted]