19 ms·
An Introduction to ZFS
- tutfbhuf 6y agoWhat are the limits of ZFS, can it scale up to the petabyte range like Ceph?
- gbrown_ 6y agoCeph is a distributed storage system whereas ZFS is not so you can't really compare the two by the capacity people run them at.
- tutfbhuf 6y agoOkay, if I'm not allowed to compare them directly, what are their different use cases?
- hosteur 6y agoI believe the z in zfs is for zettabyte
- Tuna-Fish 6y agoZFS can scale a lot further than Ceph. The limits were designed to be large enough to be never encountered in practice. Not just large enough to be never encountered by the people working on it, but large enough that it is not possible to fit a filesystem that needs more on earth, no matter how good your technology is. The only limits that can be reached is that the maximum size of a single file is 2^64 bytes, or 16 exabytes, and the maximum amount of files in a single directory is 2^48, or 281 trillion. The other limits are large enough that to reach them would generally require more energy than it would take to literally boil the oceans.
- tutfbhuf 6y agoI mean practically speaking. Lets say I actually were to buy 100 x 10TB drives. I know that I can span a ceph cluster across those drives (e.g. CERN has created some 30PB cluster in the past). Can I use 100 drives in a RAID-Z? I'm asking because the main advantage I can see for a zettabyte file system (ZFS) is at larger scale, thus I'm interested in how to actually use it e.g. for a server setup of a company that needs to save data on many disks.
- gh02t 6y agoZFS can use that many drives, but it's not a distributed filesystem so it's intended to use within a single server. For 100 disks (or some arbitrarily large number) I'd assume you want something distributed, in which case ZFS could serve as the underlying FS on each node with e.g. Gluster or Ceph. ZFS as the backing store for Gluster is fairly common, for example.
- deleted 6y ago[deleted]
- secabeen 6y agoI have two systems with 140 drives in a single ZFS pool. (14 VDEVS of 10 drives each). They work great, and we haven't had any issues with them.
- tutfbhuf 6y agoWhat type of connection do these system have (e.g. 10GbE) and how much latency? Would it be possible to span a ZFS Pool across multiple regions (us, europe, ...)?
- secabeen 6y agoThe two systems are in different buildings on campus. Each system (with 140) drives is a single unit, with 60-drive JBODs running on SAS-external connections. If you want to span a pool across multiple regions, you'll want a distributed file system on top of your ZFS to manage that. Something like Lustre or Ceph. It would still be very challenging, though.
- magicalhippo 6y ago> Can I use 100 drives in a RAID-Z AFAIK yes you could, but you wouldn't want to put them in a single RAID-Z(1,2 or 3) VDEV, but rather multiple VDEVs. This is because I/O operations scale with the number of VDEVs rather than number of disks[1]. So there's a tradeoff between space efficiency and IOPS. But AFAIK you absolutely could put 100 disks in a single RAID-Z VDEV. On the mailing lists there's frequently people posting with that or more disks in a single pool (split across multiple VDEVs). [1]: https://www.delphix.com/blog/delphix-engineering/zfs-raidz-stripe-width-or-how-i-learned-stop-worrying-and-love-raidz https://www.delphix.com/blog/delphix-engineering/zfs-raidz-s...
- rsync 6y ago"What are the limits of ZFS, can it scale up to the petabyte range like Ceph?" We (rsync.net) have scaled zpools to petabyte range. A current, example configuration would be: - 60-drive JBODs - 15 drive raidz3 vdevs, four per JBOD - 16TB SAS drives That ends up being ~192TB per vdev, 768 TB per JBOD ... and if you span a pool across two JBODs, you have ~1.5 PB. I should note that what makes it possible to sleep at night with such a configuration is the fact that raidz3 exists. If not for that, I would not configure 15 drive vdevs with jut raidz2 ("raid6") protection.
- tutfbhuf 6y agoHas rsync.net ever considered Ceph and if yes, why decided against it in favour of ZFS?
- rsync 6y agoNo, we did not ever consider ceph. rsync.net architecture is purely FreeBSD and there was just a very good fit - and roadmap - from UFS2 which we used from 2001-2012 and ZFS which we have used since. Also, for better or worse, rsync.net is about UNIX and ZFS is very unixy. We think in filesystems and files and directories and ZFS let's us keep that set of abstractions.
- ed25519FUUU 6y agoI’m in the process right now of setting up a raidz2 with 6 4TB disks. So far I’ve really liked zfs, but I have had some weird issues, like deletes taking a very long time. My plan is to upgrade the host to Ubuntu 20.04 to see if the newer ZoL version helps. Also looking forward to root volume on zfs so I can do snapshotting to the raidz2 and just let the drive fail eventually.
- imperialdrive 6y agoUhh, wow. I can barely imagine a beginner fully grasping this. I can easily imagine a professional learning from this, and it contains specialist-level insight and remarks. Great content. First time visiting STH and it's already earned my only bookmark of 2020! Thank you, Nick.
- lysp 6y agoThey have lots of great content on their news sections. As well as beginner articles, they have reviews of new upcoming enterprise cpus and hardware. Their forum is also great for information about running ex-enterprise gear at home.
- scottisbrave84 6y agoI want to do 7nm power ibm data center with zfs as only real function but no one will fund me because I hit a cop for stepping the fuck up or some shit they say they got a video of. Everyone should look and punish them for me.
- _ink_ 6y agoI love ZFS and I want to migrate all my data storage to it. The one thing that is holding me back is the inability to grow raid z with the demand. There is a pull request ongoing for a while now, but it is hard to tell when and if this feature will be available [1]. Anyone has any more insights if there is an ETA for this feature? [1] - https://github.com/openzfs/zfs/pull/8853 https://github.com/openzfs/zfs/pull/8853
- lmm 6y agoAs a home user of ZFS I would give it an ETA of never. We've been waiting for probably 10 years for this feature, but it's just never been enough of a priority from more serious (paying) users to get worked on.
- btgeekboy 6y agoThat's pretty much why I went with LVM + ext4 for my home server. Every once in a while, when I need more space, I toss another SSD in and expand the volume. Easy peasy. Though my storage needs aren't huge; I'm only up to 4TB of total SSD space.
- noja 6y agoYou can grow the pool by adding pairs of disks, if you can stomach the cost.
- redflame7 6y agoJust add another VDev. For example lets say you have 32 drives. It makes more sense to structure it as 4 x 8drive raidz2 instances rather than 2 x 16drive raidz3. Because then you will only have to add 8 drives vs 16 to expand the pool. You could incrementally expand your ZPool by adding new vdevs of the same or larger size. You could technically make your vdevs 4 drives in raidz1 but id recommend at least 6 drives per vdev in raidz2 or greater
- minimaul 6y agoI don't expect this to land any time soon, if ever. It's a very large change, and very complicated. It's been in progress ever since I started using zfs, and that was years ago :) It's worth noting you can expand a RAIDZ through replacing disks - if you start off with a pool of 4x2TB for instance say giving 6TB usable, you can expand it by replacing those disks one by one with 4TB disks - in which case you eventually end up with 12TB usable, once all disks are replaced. Alternatively, you can add another RAIDZ to the same pool with extra disks (but you will lose more capacity this way). Otherwise, recreate the pool, and restore from your backup (which you definitely have, right?). Assuming both your live and backup are zfs, this is easy with zfs send | zfs receive.
- tomas789 6y agoFor years I’m running home server. Mostly for storage plus some side projects. I only had a little time for server maintenance and I ended up redoing the whole server every time a disk failed. I was using just regular desktop grade disks so that happened every 18 months or so. This all changed when I started using ZFS. Not only it has support for raidz and mirroring (which I could get with LVM too) but it is tweekable and tunable easily. Plus commands like _zpool status_ will give you a great overview of the health of the array in no time. It might seem like nothing but it makes all the difference for me. I can recommand for everybody running a (home) server. It will save you lots of time.
- emilburzo 6y agoNot to take away from the merits of ZFS, but this sounds like something that 2 disks with software RAID1 would prevent from happening. Also, 18 months sounds kind of short, I usually get around 5 years of use from hard drives. What brand are you using? I highly recommend checking out the quarterly Backblaze HDD failure reports. Source: been running a server (read: consumer headless desktop) for years without issues.
- redflame7 6y agoMore than half the time in my experience it’s usually a failing sata cable or power-supply if the hard drive continues to functions on boot but eventually goes to degraded due to checksum/read errors. Sometimes the data is just fine but a read error occurs between the drive and controller which ZFS interprets as a failng disk
- 867-5309 6y agomodern drives have onboard diagnostics which can report errors directly to the OS. I'd be very surprised if a modern file system was unaware of this. I'd also be surprised if users of alternative file systems didn't test a suspected failing drive before concluding it was faulty
- 6y ago
- verroq 6y agoThe only thing preventing me from adopting zfs is that it is not being part of the linux kernel.
- lmm 6y agoZFS was one of the reasons I switched to FreeBSD, and it turned out to have a lot of other advantages too. Might be worth a look?
- iso947 6y agoWhen ZFS came out I tried OpenSolaris on a 48 disk machine. Regretted it immensely, that was the last sun hardware we bought Simply it didn’t fit with the hundreds of other linux boxes we had. We ended up scrapping the zfs idea and bout 1000+ disks (about 2PB) worth of linux storage over the next decade, on xfs and ext4. Had Sun’s x4500 platform worked with a Debian based linux we’d have bought that instead of supermicro. Sun lost because of their choice to exclude zfs from linux. From what I can tell, zfs is/was great, but wasn’t good enough to change our standard OS, and since then people moved to object storage
- jjav 6y ago> When ZFS came out I tried OpenSolaris Historical context: ZFS predates OpenSolaris by quite a few years, so the above statement can't be technically true.
- iso947 6y agoYes you’re correct it was normal Solaris 10, this was about 12 years ago. On my evaluation I was comparing with a satabeast and hp320s at the time, and I said that “zfs probably outweighs the problems of running Solaris” The box was still in use in 2012, we were having issues and a “zfs upgrade” was suggested, but at that point we were adding new storage on linux.
- rleigh 6y agoI did the same, and no regrets. ZFS and FreeBSD are a great combination. Wish I'd made the move several years earlier.
- hardwaresofton 6y agoSo has anyone here used zfs and btrfs and would like to comment on the differences? I've been on a heavy zfs kick lately, but the performance loss is hard to stomach and the only research I found points to btrfs being faster (though of course they both take a hit). Basically the reasons I'm drawn to zfs are: - checksumming & self-healing - ergonomics & flexibility of managing pools with zfs - copy on write for cheap local copy/experimentation (i.e. just clone your DB folder and you have a new DB) - zfs send/recv for very efficient incremental backups From what I can find it seems like btrfs does all that, and faster[0]. In addition to being faster, it also is in-kernel[1], and more flexible for the user in various ways, for example allowing resizing[2]. Looking around btrfs may not be blessed as stable but there are a lot of big orgs using it. All that said, there are articles like this one[3] which are somewhat dated but paint ZFS really positively from a maintenance point of view. Very hard to pick between these two. I'd really like to use ZFS -- the community seems very welcoming and amazing but I'm a little worried about picking the wrong tool for the job. [EDIT] - there's also this old comparison from phoronix[1] which is confusing. I'm still learning towards ZFS but sure would like to hear some strong opinions if anyone has em. [0]: https://www.diva-portal.org/smash/get/diva2:822493/FULLTEXT01.pdf https://www.diva-portal.org/smash/get/diva2:822493/FULLTEXT0... [1]: https://btrfs.wiki.kernel.org/index.php/FAQ#Is_btrfs_stable.3F https://btrfs.wiki.kernel.org/index.php/FAQ#Is_btrfs_stable.... [2]: https://markmcb.com/2020/01/07/five-years-of-btrfs/ https://markmcb.com/2020/01/07/five-years-of-btrfs/ [3]: https://rudd-o.com/linux-and-free-software/ways-in-which-zfs-is-better-than-btrfs https://rudd-o.com/linux-and-free-software/ways-in-which-zfs... [4]: https://www.phoronix.com/scan.php?page=article&item=freebsd-12-zfs&num=2 https://www.phoronix.com/scan.php?page=article&item=freebsd-...
- accelbred 6y agoI use btrfs, and the main reason I don't consider zfs instead is that zfs doesn't use the regular linux fs page cache. That and the fact that I use latest mainline kernel, so dont want to have to deal with kernel updates breaking zfs or tanking its performance, as when it lost access to the simd functions. The main feature I would want from zfs would be tiering, which can be gotten from btrfs on bcache, which is what I will likely use in the future. I think there was some issue with adding disks in zfs too? Don't exactly remember about that.
- nuker 6y agoInteresting technical comparison table https://www.snapraid.it/compare https://www.snapraid.it/compare
- nuker 6y agoI tried FreeNAS but even without data disks it ate 2.5GB ram. I went with snapraid, mergerfs and OMV combo.
- louwrentius 6y agoZFS is an awesome filesystem, this article gives a good overview. Very interesting for home users. I run it on my NAS. It is very important to realize that you can’t expand VDEVS. You can only add VDEVS. This makes expanding storage less flexible than regular MDADM RAID. Going for Mirrors is the most flexible but you lose 50% of capacity. You also get the random IOPs performance of a single drive per VDEV. Sequential performance does scale within a VDEV. You scale random I/O performance by adding VDEVS. For home usage, as a NAS, you don’t need to use SSDs for a SLOG, unless you have write-intensive random I/O workloads.
- bonestamp2 6y agoWhich OS are you using for your NAS? I ruled out Unraid and I'm looking at TrueNAS Core (FreeNAS). I want it mostly for scalable backup storage where I can easily add/remove drives from the array as I need more storage. Any others that I should consider?
- louwrentius 6y agoI am using Debian with ZoL, but it's now ancient. Just ZFS + NFS & SMB. https://louwrentius.com/71-tib-diy-nas-based-on-zfs-on-linux.html https://louwrentius.com/71-tib-diy-nas-based-on-zfs-on-linux... I would now go for Ubuntu + ZoL myself. I bought all capacity up front and paid the ZFS tax. So I don't need to expand as I go. But if you want to expand as you go, Linux + MDADM are still fine in my opinion. It's a tradeoff. Do you want to 'pay' the ZFS tax and expand at a cost, or do you want a bit more risk (on paper) but more flexibility? I ran a RAID6 of 20 drives before that using Linux + MDADM and that worked fine. And MDADM allows you to expand as you go, exactly as you want. I think the risks ZFS protect against are very small. https://louwrentius.com/what-home-nas-builders-should-understand-about-silent-data-corruption.html https://louwrentius.com/what-home-nas-builders-should-unders...
- tinco 6y agoUnraid is the only one that allows you to freely add/remove drives. And only if you're OK with taking cluster down for a minute as you add it. If you true freedom you'd need to go Ceph(FS) based, but there's no GUI to manage a Ceph cluster that I'm aware of, and they're so focused on cloud usage that you'll find little guidance on single-node clusters. edit: Forgot to read Louwrentius' comment, of course MDADM with raid 6 allows you to add disks in pairs of two, so that's definitely also an option. Don't forget to use LVM, it can get real awkward without it.
- anderspitman 6y agoI've had a FreeNAS VM set up with a couple mirrored ZFS disks (PCI passthrough) for the past few years. It's worked fairly well, but overall it's a pain to always have make sure the VM is on, mount the NFS, etc. My original motivation was preventing bitrot, but I've since seen that rationale called into question[0]. Is there a current consensus on whether ZFS is worth the hassle? I've never used the snapshotting, though I've heard good things about it. I'd much rather manage a simple local disk with offset backups. [0]: https://www.jodybruchon.com/2017/03/07/zfs-wont-save-you-fancy-filesystem-fanatics-need-to-get-a-clue-about-bit-rot-and-raid-5/ https://www.jodybruchon.com/2017/03/07/zfs-wont-save-you-fan...
- fuzzy2 6y agoWhy rely on some consensus? Instead, look at what exactly you need and how to get it. ZFS is very nice, not because of some checksums or whatever but because it offers holistic storage management including incremental backups and snapshots and compression. However, you could also use MDRaid and LVM2 Thin Volumes with any file system you like to get almost the same, without additional kernel modules.
- masklinn 6y agoThat essay seems to be essentially “hardware already does it”. If you go and read the original ZFS paper, much of the reason it exists is they found hardware lies its ass of. Hardware does not lie less since them, it lies more (witness WD recently caught essentially laying about Reds). And SMART tells you that a disk is dying, it doesn’t tell you the a disk is not dying. Furthermore the disk health tools can’t catch e.g. a dying cable.
- bleepblorp 6y agoSMART does report transfer errors between a disk and the motherboard/HBA. Failing cables should be detected.
- masklinn 6y agoThat's not correct. Some SMART implementation can report E2E errors.
- philsnow 6y ago> 256,000,000,000 / 128,000 * 70 = 140,000,000 bytes > This would be a pretty common configuration choice for a lower-end VM storage box. If you only had 16GB or of RAM in your system, all of your ARC space would be wasted with L2ARC mappings and you would only have 2GB of the entire rest of your system. Am I misunderstanding this? 140,000,000 bytes is 140MB (salesman MB, not 2^20 bytes). It looks like they're saying it's 14GB.
- Erlich_Bachman 6y ago> salesman MB There is a de-facto term distinction: MB megabyte (1000^2) vs MiB Mebibyte (1024^2). https://en.wikipedia.org/wiki/Mebibyte https://en.wikipedia.org/wiki/Mebibyte
- npteljes 6y agoWhile planning my ZFS setup back then, I had good fun with this ZFS capacity calculator. The numbers turned out to be a little different but I still think it's worth a try. https://wintelguy.com/zfs-calc.pl https://wintelguy.com/zfs-calc.pl
- olavgg 6y agoOpenZFS 2.0 will come with two awesome features that I am really looking forward to. The first one is ZSTD compression, this will work great together with MySQL and PostgreSQL. The second major feature is persistent L2ARC, I'm using LARGE ssd's as cache. And the warmup time takes weeks. So rebooting has a major performance impact. For the last 10 years, I have been using FreeBSD with ZFS. This has been working perfectly with good performance. But now I want to take advantage of even faster network speeds, with RoCE/RDMA. And FreeBSD support for iSER, NVMe-OF is non-existant, while Linux has excellent support for these technologies.
- nix23 6y ago>iSER https://porter.io/github.com/sagigrimberg/iser-freebsd https://porter.io/github.com/sagigrimberg/iser-freebsd >NVMe-OF https://papers.freebsd.org/2018/bsdcan/shwartsman-roce_as_a_performance_accelerator/ https://papers.freebsd.org/2018/bsdcan/shwartsman-roce_as_a_...
- trasz 6y agoSupport for iSER (only initiator side, sadly) has been merged years ago: https://www.freebsd.org/cgi/man.cgi?iser https://www.freebsd.org/cgi/man.cgi?iser.
- MaXtreeM 6y agoHas anyone tried to do a home NAS server with ZFS on Raspberry Pi 4?
- jasomill 6y agoWhy would you choose a system with a single Gen 2 PCIe lane xor bandwidth-constrained USB 3.0 as the only high-bandwidth I/O to build a storage server? I ask as an owner and regular user of several Pi models, including a 4. With that said, assuming you choose PCIe, acquire a decent HBA, and build it into a decent enclosure, I'm sure it'd work as well as typical low-end off-the-shelf NAS boxes, and wouldn't cost that much more in time and materials to set up.
- gpanders 6y agoFor the uninitiated among us (including myself) could you explain what an HBA is? Or share any more details on how you’d build a low-cost NAS that’s not a Raspberry Pi? I have a Raspberry Pi 4 that mounts a USB hard disk and serves files over SMB and Nextcloud. I have been considering reformatting the drive to ZFS or btrfs and booting the Pi directly from that so that I can start taking snapshots. Is this a bad idea? I’ve looked at buying dedicated NAS hardware before (mostly Synology products) but I’m always deterred by the cost. A low end Synology NAS with drives runs around $500 or $600, which is a huge jump from my little Pi.
- magicalhippo 6y agoHBA is basically a PCIe to SAS/SATA card (could be other variants but this is the typical one). For a low-cost NAS you could use a spare computer. My first NAS was my old desktop computer, using the motherboard SATA ports. That said, ZFS should work on the Raspberry Pi 4 at least in 64bit mode. You probably will want to use Ubuntu as it has ZFS support. If you have a spare drive you can test with, give a whirl. Just keep in mind this[1] if you have poor performance from the USB-SATA. [1]: https://www.raspberrypi.org/forums/viewtopic.php?t=245931 https://www.raspberrypi.org/forums/viewtopic.php?t=245931
- netflixandkill 6y ago
- polskibus 6y agoCould a ZFS server farm be a good alternative to a smallish Ceph installation ? If I wanted to build a storage layer for a small private DC - is Ceph the way to go? ZFS? Or maybe there's something else? I'd like to have storage details hidden away from storage layer so that applications can write to a mount or something similar - what is the best solution for such problem these days, if one needs resiliency and redundancy baked in ?
- HeadsUpHigh 6y agoZfs sounds like a good solution although I'm not experienced with Ceph. Btrfs is ab option too but it has sone pitfalls.
- sekh60 6y agoI use Ceph in my homelab. Small cluster of five all-in-one noses with ten OSDs each. I went this way since it gives me a lot more expandability and fault tolerance than ZFS can. I use min_size 2, size 3. My use case is both CephFS and RBD for a small two node OpenStack cluster. I have found CephFS to be rather performant for my use cases, enough that I had to get a 10Gbps switch. I am not machine or the bandwidth on that switch, but my individual clients use more than 1Gbps. OpenStack and Ceph tie together wonderfully. I have my VMs backed by NVMe drives and my VMs are snappy. Recovery is quick too. I am using crappy first gen xeon-d boards and even with those I hit 8Gbps recovery on those drives. Ceph shines when you have a lot of parallel access. It is recommended to have at least ten nodes for a production cluster so recoveries so not take too long. If you have a lot of clients Ceph is king. I used ZFS in the past as a simple Fileserver before using Ceph. It worked well, I could saturate a 1Gbps link, however I found the vdev resize limitation too restricting at times when I wanted to expand by a little bit. It is pretty easy to manage, though I find Ceph very easy to manage as well. For my backup server which is a target for BorgBackup I went with btrfs for the better flexibility it offers with resizing arrays.
- azalemeth 6y agoHow many computers do you have at home, in your homelab? Do they also heat your house? Where do you store them?
- qatanah 6y agoLoving how zfs is so easy once you get a hang of it. I'm trying to test it out on an i3a.large instance with 1.2TB NVME ssd and benchmarking postgres on it. Trying to move out of RDS since time to time I have an heavy iops scripts. Only thing left is doing zfs snapshots next for my backups. zfs snapshots or pg_dump snapshots? I wonder what's better.
- tristor 6y agoI've been using ZFS in production both in my home and at work since 2012. It's come a long way in FreeBSD, and I think is now quite clearly the best filesystem choice for nearly every workload. ZFS on Root is super usable and easy to install now, and is great with an SSD mirror. Highly recommended as something to learn and use, it's obviously the best choice for a home built filer, but is also an excellent choice even for general purpose server use. My colo setup is running FreeBSD w/ ZFS which is very stable for backending VPN servers, web servers and app servers of all stripes, etc.
- earthscienceman 6y agoYou know, these articles always come up in the context of fileservers but... ... for me using ZFS has changed the way I look at files, filesystems, data, and backups for general computing. I've been a linux user for 13 years but never felt the need to have a fileserver. Now being able to plug a drive in and take a snapshot without rsyncing or thinking about what I'm snapshotting, having it be inherent to the filesystem, was a game changer. Not to mention being able to snapshot important folders to the native drive in case I need to recover a file from a previous state. I run datasets for categories of data and I can choose categories that I want regular local snapshots of (zvol/crypt/Documents, zvol/crypt/scripts, zvol/crypt/Papers) Essentially, ZFS manages my files for me. And it all comes with things I didn't know I needed, like filesystem compression. I know BTRFS also attempts to provide this, and there's the licensing issues with ZFS, but I wanted MacOS compatibility also. Although that was an adventure on its own.
- matheusmoreira 6y ago> Essentially, ZFS manages my files for me. Yes! Managing our files is the whole point of file systems! It's amazing how bad at it most of them are. Linux is still catching up with btrfs... It's extremely aggravating how most file systems can't create a pool of storage out of many drives. We end up having to manually keep track of which drives have which sets of files, something that the file system itself should be doing. Expanding storage capacity quickly results in a mess of many drives with many file systems... Unlike traditional RAID and ZFS, btrfs allows adding arbitrary drives of any make, model and capacity to the pool and it's awesome... But there's still no proper RAID 5/6 style parity support.
- earthscienceman 6y agoIt really is incredible. ZFS datasets are essentially data collections, and I can categorize my files according to how I want them managed. It's essentially what iOS dreams of but in a much more manageable, configurable, and open way.
- sly010 6y ago
- pstch 6y agoI find it very unfortunate that file metadata is not encrypted : if you need this to be encrypted, you need to stack LUKS on top of ZFS, and you lose many of the benefits of ZFS (per-dataset encryption, healing ability, RAIDz, etc) while doing so. Running ZFS->LUKS->ZFS to recover some of these benefits is also not feasible at all (ZFS doesn't like to self-host, even through a virtual machine).
- nichch 6y agoWhat file metadata is unencrypted?
- pstch 6y agoI was mistaken, file metadata is indeed encrypted.
- hikarudo 6y agoWouldn't just LUKS->ZFS be enough?
- robbyt 6y agoWhy does metadata encryption matter?
- Nokinside 6y agoZFS encrypts most metadata. Metadata not encrypted: Dataset / snapshot names, Dataset properties, Pool layout, ZFS Structure, Dedup tables ZFS encrypts: File data and metadata ,ACLs, names, permissions, attrs Directory listings,, All Zvol data,FUID Mappings ,Master encryption keys ,All of the above in the L2ARC ,All of the above in the ZIL For most uses and use cases this is net increase in security. You can do some operations on data without needing the keys.
- pstch 6y agoOh it seems I was mistaken about that. ZFS does encrypt enough metadata indeed. Sorry for the noise.
- GekkePrutser 6y ago> "Use CMR with ZFS, not SMR"... Well, yeah, good point to make but it is becoming harder and harder to find CMR drives, especially in 2.5". ZFS should really adapt to this. Perhaps using bigger block sizes or something. Because SMR is not going away.
- netflixandkill 6y agoUpping the block size to more than whatever SMR overlaps on the drives works for workflows that don't care about consistent random rewrite performance. Can get 120 MB/s sequential writes on the cheap seagate SMR drives in my big slow pool which is enough for clients on gigabit or less networks. Although frankly at this point if you care about performance you're on NVME anyway.
- phil21 6y ago"No SMR with ZFS" is probably a bit overblown. It simply depends on your use case. Rebuilds certainly are a problem but beyond that I think it's simply "be ok with slow drives". I've been operating a decently sized SMR pool for 3 years now with no major issues, including surviving two drive failures. If you try to put random write workloads onto SMR you're gonna have a bad time no matter what you do. In my use case it's great, since this is effectively WORM storage writing giant files to disk all at once so I have effectively 0% fragmentation and all subsequent reads of the file tend to be sequential. That said, this is for my personal use and lab projects. Not sure I'd go ZFS+SMR for a production workload.
- vermaden 6y agoThat depends what is your use case. I use two 5TB 2.5 SMR drives in ZFS mirror: https://vermaden.wordpress.com/2019/04/03/silent-fanless-freebsd-server-redundant-backup/ https://vermaden.wordpress.com/2019/04/03/silent-fanless-fre... These drives can slow down to 30-40 MB/s when filled to 80% or more but I use that storage over WiFi which is at most 11-12MB/s which means the SMR problem does not exist for me. If I would be using that storage over LAN the 30-40 MB/s in 'WORST CASE' is also not bad considering that maximum real life LAN speed over gigabit network is about 80-90 MB/s. Its also not possible to get non-SMR large 2.5 drives. I use 2.5 drives as they are silent and they need very small amount of power comparing to 3.5 drives.
- zepearl 6y agoAny remarks/experiences about the "record size" of ZFS, maybe especially in relation to RAIDZx? I don't fully understand it. I have in a NAS a RAIDZ1 of 4HDDs on which I set a recordsize of 1MB, currently full at 50% and so far performance has been good with both big and small files... . I'll probably create in a future a RAIDZ2/3 by using ~8 HDDs and I'll test various record sizes but I just wanted to know if anybody had already any positive/negative experiences with some combination of record size and RAIDZx... . Thx
- seized 6y agoThe record size setting is the max, ZFS can use less in some cases. The best one depends on the data youre writing and the ashift of the pool, so testing is best. Large record sizes are helpful for large files (less overhead losses). Some good beginner info https://arstechnica.com/gadgets/2020/05/zfs-versus-raid-eight-ironwolf-disks-two-filesystems-one-winner/ https://arstechnica.com/gadgets/2020/05/zfs-versus-raid-eigh... Some advanced info https://www.joyent.com/blog/bruning-questions-zfs-record-size https://www.joyent.com/blog/bruning-questions-zfs-record-siz...
- zepearl 6y agoThank you! I did read Arstechnica's article in the past but I did not feel comfortable with their results... (I'm not challenging them, I'm just not sure if they're relevant for me or not). So, I just did a test (ashift 12, RAIDZ1 with 4 8TB HDDs) and I got better performance in both cases with a 1MB recordsize vs. 128KB (all sequential). Recordsize 1MB: reading one 10GB file: 21 seconds. reading 10000 1MB files: 83 seconds Recordsize 128KB: reading one 10GB file: 31 seconds reading 10000 1MB files: 116 seconds Maybe a small recordsize can have some benefits when overwriting parts of the files...mmmhhh...? Ok, it seems complicated => I'll just have to test different variants :)
- radiowave 6y ago> Maybe a small recordsize can have some benefits when overwriting parts of the files...mmmhhh...? Right. People who've done more testing than me reckon on 16KB being a good record size for transaction-processing database work, where tables are seeing lots of small inserts and updates. (You might think matching the database's block size would be ideal, e.g. Postgres writes 8KB at a time, but the rationale here is that you tend to get better compression at 16KB recordsize than 8KB, and the benefit from this outweights the write-amplification.) But if database update performance isn't a big deal for you then you can probably just ignore this. I've not done any testing of my own at the 1MB size, but I don't think I'd be inclined to try it unless I was fairly confident that there weren't going to be many small writes to big files. In short: use the large recordsize where you think you've got a good case for it, and likewise with a small record size. Otherwise, just stick with the default.
- gen220 6y agoI saw this deck on ZFS and btrfs while researching them some years ago [1]. It led me to think ZFS would be fine for me to run in my closet/personal cloud, but I’d probably want to avoid it at work. I know it’s an old deck... has anything substantial changed since it was put together? Specifically around the licensing [2], inclusion in the kernel, etc. It looks like FB in particular had invested a lot in btrfs [3] over ZFS. [1]: http://marc.merlins.org/linux/talks/Btrfs-LCA2015/Btrfs.pdf http://marc.merlins.org/linux/talks/Btrfs-LCA2015/Btrfs.pdf [2]: https://arstechnica.com/gadgets/2009/10/apple-abandons-zfs-on-mac-os-x-project-over-licensing-issues/ https://arstechnica.com/gadgets/2009/10/apple-abandons-zfs-o... [3]: https://btrfs.wiki.kernel.org/index.php/Contributors https://btrfs.wiki.kernel.org/index.php/Contributors
- tw04 6y agoSure, I can help. That powerpoint is horribly out of date and I'd argue full of FUD (intentional or not). >• Raid 0, 1, 5, and 6 are also built in the filesystem BTRFS RAID has been a nightmare from day 1 including complete data loss. Parity-based RAID has been "just around the corner" for almost a DECADE. Sure you can do RAID-1 but in 2020 I'm just not interested in losing half of my capacity. Yes you can layer it on top of MDRAID but that eliminates half the elegance of what ZFS brought to the table. >ZFS is fairly memory hungry, it's recommended to have 16GB of RAM and give 8GB or more to ZFS (it wasn't designed to use the linux memory filesystem, so it uses its own memory that can't be shared with the rest of linux). This is just flat wrong. You need 2GB of memory for a happy filesystem, 8GB+ is if you're doing deduplication which is unnecessary overhead in most environments. I'm also not sure where the "memory that can't be shared with the rest of linux" is coming from - ARC will use and free memory as needed by the system. >Due to the CDDL being incompatible with GPLv2, a linux vendor or hardware vendor will never be able to ship a linux distribution or hardware device using ZFS Except Ubuntu already is. I believe SLES does as well. >As a result, you shouldn't plan on using ZFS for any product that you might ever want to ship one day. Delphix ships a product today based on ZFS. 0 issues. >Oracle may have stopped further work on ZFS as a result. Or it could be another reason entirely... Oracle absolutely didn't stop work on ZFS, I'm not even sure where he came up with that nonsense. Oracle continued to update and release new versions of ZFS long after the lawsuit was settled. You know what's far more telling? 13 YEARS after starting BTRFS: Oracle uses ext4 as their default filesystem, not BTRFS. Redhat has dropped support for BTRFS entirely. https://fossbytes.com/red-hat-deprecate-btrfs-filesystem-stratis/ https://fossbytes.com/red-hat-deprecate-btrfs-filesystem-str...
- chokeartist 6y agoI have been a FreeNAS (ZFS) user for years. While I do follow a proper rule of 3 for backups, my main FreeNAS volume (100TB) is where I will randomly dump stuff until I sort/backup it up later. It hosts a wide range of stuff from media, to small files (db). The only thing that will eat that data is a hacker, massive sw bug, or my house burning down. I've broke a lot of systems and data over the years. With that in mind, I like my data storage (NAS/filer) to be boring and predictable. FreeNAS is exactly that for me.
- js2 6y agoWhat are you using for H/W and what drives compromise your 100TB volume? Do you know what the idle power draw is? I'm running FreeNAS on an ancient Lenovo minitower with just four drives as two mirrored volumes and it's about time I upgrade the thing. But it just works and only draws about 60W.
- chokeartist 6y ago30x 4TB HGST drives is the array. Running on multiple LSI HBAs with a NVMe drive for write cache. Idle power draw is 250watts-ish including the host hardware (cpu/mem/net). I think I will go cry now that I am thinking of my power bill.
- whitepoplar 6y agoHow should one best protect a home NAS/fileserver from being hacked? I'm in the process of putting one together, but this one makes me nervous. Is it simply having an up-to-date OS, setting up, say, an SSH bastion/wireguard for remote access, and calling it a day?
- chokeartist 6y agoSorry I missed this one - my bad. I do a few things including A) keep it patched B) minimize the number of exposed services C) utilize pfsense to control [network-level] access above and beyond what freenas itself provides D) stream all host/service logs to an ELK stack to review for any funny business that may occur. Not bulletproof but I haven't had an incident yet. Hope this helps! P.S. All of this assumes proper backups. I can restore most of my stuff from backups, it just will take forever.
- roflc0ptic 6y agoTangentially related to ZFS, I’d like some advice. I’ve got a couple of projects I’d like to do that require maybe 100TB of storage: some scientometrics against the sci-hub collection, as well as building a bajillion scala projects from github. I don’t really care about data redundancy. The cheapest way I’ve been able to figure out to do this is just buy a case with 15 HDD bays, eg the Anidees AI crystal case, and just get a ridiculously beefy processor + 256gb RAM and do all the computation on a single box. Does this sound right? All of the purchasable NAS cases all seem more expensive, but I’m out of my element. I expect I’m going to want to figure out ZFS to make a single logical drive. Does this seem right? Building eg a backblaze pod is outside of my budget, and my eyes glaze over whenever I try to read about NAS controllers.
- threatripper 6y agoThis sounds about right. Not very reliable but cheap and easy to set up. If you make a software raid (e.g. with lvm) you could as well use any other file system.
- secabeen 6y agoYep. If you're looking for more specific recommendations, the community at /r/datahoarder is a great source. With shucked USB drives, and just the biggest ATX case you can find, your costs should be pretty low.
- roflc0ptic 6y agoThanks for the input, both of y’all. Roger that about shucking drives. It is strange to me that ~10 TB external drives end up being the cheapest solution, but happy to capitalize on it
- jcastro 6y agoThis is what I did, I just snagged some cheap Rosewill case off of newegg with 12 bays, a PCIe sata card and packed the case with drives. You can slice and dice the drives however you like. I found this guy's blog posts to be useful for running a homegrown NAS: - https://louwrentius.com/should-i-use-zfs-for-my-home-nas.html https://louwrentius.com/should-i-use-zfs-for-my-home-nas.htm... - https://louwrentius.com/the-hidden-cost-of-using-zfs-for-your-home-nas.html https://louwrentius.com/the-hidden-cost-of-using-zfs-for-you...
- rsync 6y agoRelated article from arstechnica: https://arstechnica.com/information-technology/2015/12/rsync-net-zfs-replication-to-the-cloud-is-finally-here-and-its-fast/ https://arstechnica.com/information-technology/2015/12/rsync... You can issue arbitrary 'zfs send' to rsync.net, over SSH.
- ed25519FUUU 6y agoMy main gripe with ZFS is that I can't dynamically size up an array once it's created. A raidz2 can't go from 6 disks to 7 or 8 disks without completely destroying and recreating the array.