8 ms·
> ZFS never really adapted to today’s world of widely-available flash storage: Although flash can be used to support the ZIL and L2ARC caches, these are of dubi
by floatboth 9y ago
> ZFS never really adapted to today’s world of widely-available flash storage: Although flash can be used to support the ZIL and L2ARC caches, these are of dubious value in a system with sufficient RAM, and ZFS has no true hybrid storage capability.
How is L2ARC not "true hybrid"?
> And no one is talking about NVMe even though it’s everywhere in performance PC’s.
Why should a filesystem care about NVMe? It's a different layer. ZFS generally doesn't care if it's IDE, SATA, NVMe or a microSD card.
> can be a pain to use (except in FreeBSD, Solaris, and purpose-built appliances)
I think it's just a package install away on many Linux distros? Also installable on macOS — I had a ZFS USB disk I shared between Mac and FreeBSD.
Also it's interesting that these two sentences appear in the same article:
> best level of data protection in a small office/home office (SOHO) environment.
> It’s laughable that the ZFS documentation obsesses over a few GB of SLC flash when multi-TB 3D NAND drives are on the market
Who has enough money to get a mutli-TB SSD for SOHO?!
- edude03 9y ago> How is L2ARC not "true hybrid"? L2ARC in my understanding is only for reads, whereas ZIL is the write ahead log. Ideally, ZFS would "combine" the two into a MRU "write through cache" such that data is written to the SSD first, then asynchronously written to the disk after (ZIL does this already) but then, when the data is read back, it's read back from the SSD.
- philsnow 9y agoThis makes sense; if you were building a service that served bytes off of disk, with an in-memory LRU cache, you wouldn't have two separate pools of memory, with one for writes and one for reads. Did the ZIL and L2ARC concepts come up before SSD was widely available? Especially the ZIL seems very much optimized for crazy enterprise 15k rpm spinning rust. Memory and SSD access characteristics are so different from spinning disks; I don't know why ZFS separates ZIL and L2ARC.
- koffiezet 9y agoBecause the ZIL is a journal log, not a cache. It is intended to increase data security without sacrificing too much performance. Many people also confuse ZIL and SLOG devices though... By default, the ZIL is written on the same disks as where the data will be stored, but an external device (aka SLOG) can be added. From that point on, your write IOPS will be limited by this SLOG device, so normally you add a more expensive fast disk as SLOG device to increase your write IOPS.
- notacoward 9y ago> Did the ZIL and L2ARC concepts come up before SSD was widely available? Yes. After they realized that their initial claims about not needing such things was bullshit (which some of us had told them at the time) but before SSDs became common.
- ori_b 9y ago> How is L2ARC not "true hybrid"? As far as I'm aware, it's not persistent. Reboot, and your cache of recently accessed files is gone.
- Borealid 9y agoA 2TB SSD costs around 400-500USD. That's not exactly out of the realm of possibility for a small office.
- floatboth 9y agoSure, but it's a rather high price for just faster storage. And like… 2 TB of cache is a bit high for SOHO NAS, and if you go full SSD for storage, you'd want two of them for a mirror and that's 1000 USD already…
- loeg 9y agoProbably better to spend $150 on SSD cache and $150 on RAM cache than to spend $300 on RAM and $0 on SSD, though. Or the other extreme.
- dragontamer 9y agoAnd we spend what? $200 upgrading from an i3 to an i7 for like... 30% faster CPU performance? Upgrading $100 hard drive to $400 SSD results in like, 500% improvements in storage speed. If you have any storage-related task... such as video editing, handling of large datasets and whatnot... the SSD will have a far bigger impact on your productivity than any CPU upgrade.
- ioquatix 9y agoI've got 6x 4TB WD RED in a ZFS RAID10 (3x2 mirrors). I get about 500Mbytes/s read and write, on average. You start getting into 10GbE territory pretty easily with even consumer drives and ZFS. With SSDs, you'd quickly need trunked 10GbE if you want to fully saturate your network - in addition to that consider the client requirements - e.g. each workstation with 10GbE or 1GbE. You'd have to have some pretty decent requirements to necessitate a ZFS RAID10 with [NVMe] SSDs.
- hvidgaard 9y agoThe benefit of SSD is not the raw sequential transfer. That is easy to max out with spinning platters. You want SSD for random access, it's significantly better at that.
- zackelan 9y ago> How is L2ARC not "true hybrid"? I think what that may be referring to is that the ARC is in-RAM and obviously cleared on a reboot, so as a result L2ARC on an SSD is also not persistent. After a reboot, you have to allow the ARC to fill up, then as it evicts data from the L1ARC it's pushed to L2ARC. Until that happens the SSD is not used at all.
- binarycrusader 9y agoARC is not necessarily cleared on reboot (if doing a "fast" reboot; that is, kernel reload); that's platform and implementation-dependent. Lookup "Persistent L2ARC".
- cryptonector 9y agoPerhaps "true hybrid" == RAM->SSD->HDD->tape, that is, warm-through-cold storage? Whatever. What would be really nice is raw SSD/storage access so that ZFS (or other FS) could manage all the wear leveling and bad block mappings.
- nomel 9y agoWear leveling is a implementation detail of the medium that can and will change. It doesn't seem right to put that in the filesystem itself. If anything, I would think it should go into some 'generic <specific flash technology here>' device driver.
- tw04 9y ago>. It doesn't seem right to put that in the filesystem itself It absolutely belongs in the filesystem. When you're doing RAID of any type across the devices, you need that layer to manage the underlying media. A single device view will never appropriately manage wear leveling and garbage collection. There's a reason companies like NetApp have been working with drive vendors to have more control over the underlying media: http://www.samsung.com/us/labs/pdfs/2016-08-fms-multi-stream-v4.pdf http://www.samsung.com/us/labs/pdfs/2016-08-fms-multi-stream...
- hvidgaard 9y agoIt's difficult problem to solve. If we do put it in the filesystem, we need some way for the filesystem to know how the NAND behaves. Otherwise the lowest common denominator dictates and nobody gain anything. With this in mind, I don't see why the disk cannot handle this in the firmware. As long as there is enough free NAND on the drive, it can manage wear level and GC just fine assuming that it gets TRIM commands.
- tw04 9y agoYou've just described why we have storage appliances for high performance and enterprise workloads, and why the drive to standardize always swings back around to customized software and hardware.
- chongli 9y agoWhy should a filesystem care about NVMe? Because at some point the filesystem becomes a bottleneck. ZFS was designed with the assumption that CPUs would be way faster than storage. When you get speeds over 10GB/sec, [0] you are going to spend a lot of time checksumming all that data. [0] http://www.seagate.com/ca/en/about-seagate/news/seagate-demonstrates-fastest-ever-ssd-flash-drive-pr/ http://www.seagate.com/ca/en/about-seagate/news/seagate-demo...
- wmf 9y agoWhat could ZFS do differently to solve that problem while maintaining data integrity?
- loeg 9y agoJust one idea: offload checksum calculation to a DMA engine. Linux already has a generic DMA engine facility in the kernel, backed by e.g. I/OAT on some Intel hardware.
- boomboomsubban 9y agoAssuming this is the same thing as hardware-assisted checksums, both the ZFS on Linux maintainer and Intel have said that they are working on it at various cons last year.
- olavgg 9y agoFletcher checksums are very cheap and not a bottleneck. https://github.com/zfsonlinux/zfs/issues/4789#issuecomment-230764245 https://github.com/zfsonlinux/zfs/issues/4789#issuecomment-2...
- chongli 9y agoMaybe I'm reading those benchmarks wrong, but they appear to max out well under 10GB/s. This would mean you'd be CPU bound on your checksums alone with one of those Seagate cards.
- fweespeech 9y ago> Who has enough money to get a mutli-TB SSD for SOHO?! https://www.amazon.com/Crucial-MX300-Internal-Solid-State/dp/B01KKZLX46/ref=sr_1_1?s=pc&ie=UTF8&qid=1499890752&sr=1-1&keywords=2tb+ssd https://www.amazon.com/Crucial-MX300-Internal-Solid-State/dp... I was contemplating a build with 2 of these in a RAID 1 configuration for my next homelab server. Personally I run a gaming (windows) desktop at home and an always-on UPS-backed homelab server that handles minor ops tasks (mostly backing up side projects and some ETL) + provides dev VMs. My home office "budget" is ~$2k/year. My gaming/work desktop is ~4 years old, represents ~$2.5k of that budget. Monitors/peripherials/desk/chair generally eats another $2k and are of a similar age. I generally spend ~$2.5k on the dev server. I then usually toss ~$1k into a laptop. I could easily see someone who purely works from home (rather than 1-2 days a week) operating with a larger budget and genuinely needing a ZFS setup of 2TB SSDs. Realistically, $2-3k/year is 2-5% of the sort of salaries we see on HN given we make a living at this sort of thing...it isn't surprising people would spend that kind of money to me. https://hurdlr.com/blog/software-web-developer-tax-deductions https://hurdlr.com/blog/software-web-developer-tax-deduction... Keep in mind "business equipment" certainly qualifies for such a dev server so you won't be paying taxes on it if you itemize as well.
- snuxoll 9y ago> Keep in mind "business equipment" certainly qualifies for such a dev server so you won't be paying taxes on it if you itemize as well. Even factoring in my homelab spend I don't get any more itemizing than taking the standard deduction as a married individual making $85K, and I spent well close to $2000 on it last year.
- sfoskett 9y agoHybrid storage combines flash and disk for better performance and lower cost. Most hybrid storage is tiered, meaning that data can "live" on either flash or disk, but there are other approaches that look a lot more like a cache (see Nimble Storage, for example). True hybrid storage would work like Apple Fusion Drive or Hybrid Storage Spaces Direct - A single pool with SSD and HDD where data can reside wherever is best. L2ARC is only ever a cache and although it can be on SSD, ZFS generally won't use much more than a few tens of GB. ZIL isn't even a cache and is only used for synchronous writes. Check any ZFS tuning guide and the gist will be "just buy more RAM or create an all-SSD pool" rather than trying to wedge an SSD into L2ARC or ZIL.
- wfunction 9y agoHow reliable are hybrid storage drives? I feel like the last one I saw failed pretty shockingly quickly but not sure (it wasn't mine)...
- sfoskett 9y agoI'm sorry, I was referring to hybrid storage as a technology category, not the hybrid disk drives like Seagate Momentus XT. You're right that a 2-drive hybrid (Fusion Drive) will mathematically be less reliable than a non-hybrid one and that those "hybrid disks" haven't lived up to the hype. Pretty much every enterprise storage solution designed today is hybrid (SSD plus HDD) or all-flash and includes lots of advanced availability features.
- wfunction 9y agoOhh okay, gotcha.
- adrianratnapala 9y ago> "Pretty much every enterprise storage solution designed today is hybrid" By this do you mean simply that there is both SSD and disk storage around, or do you mean that the storage system transparently chooses where to store particular things without the apps having to care. Because the latter thing is not my experience.
- vbezhenar 9y ago> I think it's just a package install away on many Linux distros? Also installable on macOS — I had a ZFS USB disk I shared between Mac and FreeBSD. It's easy to install if you want to use it as an additional filesystem. But if you want to install e.g. RHEL on root ZFS, it's quite an adventure. Even Ubuntu with first-class support for ZFS does not support ZFS on root from the box. Actually I don't understand it. Of all the features, snapshots looks like killer feature for Linux distributions. Make snapshot before upgrade, allow easy rollback if upgrade gone wrong. Something like Windows restore points, but much more reliable.
- phil21 9y agoI agree, snapshots are amazing for / on a linux machine. Makes backups actually work right, and be very inexpensive (computationally). But, the tooling is still rather new and untested. At work we are still stamping out upstream bugs that really shouldn't exist, but the ecosystem is certainly getting better. It's only been a couple years since the ZFS on Linux project has gotten remotely any uptake by the distro folks, and even now that interest is rather tepid due to the licensing issues. If ZFS was GPL I think it would have been the default filesystem on Linux for quite some time now.
- floatboth 9y ago> if you want to use it as an additional filesystem Which is totally fine for a NAS! But yeah, I like my ZFS root on all the FreeBSD installs :) > snapshots looks like killer feature for Linux distributions True. FreeBSD and of course Solaris/illumos had boot environments for a long time, it's an excellent feature.
- ScottBurson 9y agoSince Btrfs also has snapshots, it seems perfectly adequate for a root FS. It hasn't seemed important to me to run a ZFS root, even though I have all my user data in ZFS.
- RJIb8RBYxzAMX9u 9y agobtrfs snapshots have a lot of gotchas. Quoting myself: https://news.ycombinator.com/item?id=14724820 https://news.ycombinator.com/item?id=14724820
- raattgift 9y ago> How is L2ARC not "true hybrid"? It doesn't persist across import/export or reboot. It is demand-filled. Not all data in the main vdevs are eligible for l2arc. There is memory overhead for l2arc buffers. There is CPU overhead in processing l2arc headers. Once a buffer is in l2arc it stays in l2arc until the underlying data is overwritten or destroyed, or until the l2arc has filled up and the buffer is replaced with fresher data. A true hybrid in the zfs context would let one pin a dataset or zvol onto a particular vdev, or pin only the (zfs) metadata (or a subset thereof) of a dataset, zvol or pool to a particular vdev. Openzfs will eventually get both persistence and this form of true hybrid. Unfortunately automatic migration by zfs of hot data to low-latency vdevs and cool data from low-latency vdevs is not really possible without solving the infamous block-pointer-rewrite problem.
- burntrelish1273 9y agoBoth the L2ARC and the ZIL can be put on separate vdevs (IIRC), to use faster or more reliable SSDs. L2ARC typically wants a pair of striped fast SSD vdevs while the ZIL should be on a mirror vdev of higher-reliability (SLC-like) SSDs. Also, be sure to have 8+ GiB system RAM available at all times or performance is gonna suck.
- Gonzih 9y ago> I think it's just a package install away on many Linux distros? Also installable on macOS — I had a ZFS USB disk I shared between Mac and FreeBSD. Having your root on a filesystem that is provided with your kernel is not ideal situation, update issues make your system basically unbootable.
- mrkgnao 9y agonot* provided?
- keeperofdakeys 9y agoFrom a protocol level, NVMe is actually designed for SSDs - just look at the amount of queues it has https://en.wikipedia.org/wiki/NVM_Express#cite_ref-ahci-nvme_5-3 https://en.wikipedia.org/wiki/NVM_Express#cite_ref-ahci-nvme.... To really take advantage of this, you'd need your filesystem to be designed for many independent IO streams. ZFS metaslabs probably help, but I'm sure there is more you can do in this area. On the other hand, I don't think many people would be hitting any limits where this matters.