4 ms·
What superior architecture does zfs have?
by tacticus 13y ago
What superior architecture does zfs have?
- rch 13y agoMy understanding is that btrfs is still in catchup mode for the foreseeable future, but might eventually cover the distance. Has btrfs jumped ahead of zfs in ways I haven't heard about? Edit - this is my first search result: http://rudd-o.com/linux-and-free-software/ways-in-which-zfs-is-better-than-btrfs http://rudd-o.com/linux-and-free-software/ways-in-which-zfs-...
- midas007 13y agoI'd be biased to agree since that's Manuel's blog, someone I used to work with. I've supported 24x7 and 9x5 ops where downtime was unacceptable. zfs makes it a whole lot easier to perform upgrades, know data and metadata are solid and send snapshots around.
- rch 13y ago> ZFS uses atomic writes and barriers This about settles the question for me. Assuming that the implication that btrfs performs otherwise holds true.
- midas007 13y agoYeah, it depends on the use case. For home directories and large risk items like financial stuff, testing that barrier writes are happening is a good thing. You don't want a storm to knock out a DC to learn that the hw/sw fs stack was lying to you at some level.
- tacticus 13y agoBarriers are also used in btrfs. and from what i can tell clone operations are also handled atomically though i kinda wonder what exactly is meant by atomic writes.
- rch 13y agoThanks for clearing that up. I might have to dig into it a bit more.
- tacticus 13y agothe first parts are right (though i don't know how accurate) and could certainly be improved with tooling. testing will come with time. with raidz sure it's great but the vdevs being immutable is really rather annoying. the way btrfs handles multi device stuff is significantly better (replication is not between 2 devices but closer to the file (it allocates a chunk of space and decides where to put the other replica in the pool). though i wish the erasure coding stuff would land faster. btrfs has had send\receive for a while. I haven't needed to dig into the btrfs man pages yet so can't common on how accurate this is. btrfs also uses barriers log devices and cache devices are awesome. i hope btrfs adds them. the block device thing is a limitation of btrfs and annoys me though i've slowly moved to just having files. (though in anything largish i would probably be moving to a distributed fs anyway) the sharing stuff is good but i think that's a tooling issue not an fs issue. btrfs has an out of band dedup allowing you to run periodic dedup without having the memory penalty of live dedup (though costing disk)
- midas007 13y agoOnline scrubbing, so no downtime waiting for fsck for one. If you'd used it, you'd know how many hard won production battles solaris devs poured into making zfs better from the ground up. btrfs is oracle's NIH syndrome, reinventing the wheel instead of developing one that already had, pun intended, traction.
- tacticus 13y agobtrfs has online scrubbing. btrfs is the response to sun picking an incompatible license. when that is removed zfs might get more interesting for a lot of people.
- midas007 13y agoOracle acquired Sun, so they could have solved it by just choose another license moving forward. One gotcha is the ZFS (Solaris core) team vehemently resisted anything GPL-compatible. Something like a BSD license would make the most commercial sense. Instead, Oracle has a consistent pattern of losing community goodwill that loses customer interest and pushes developers to fork.
- tacticus 13y agoiirc their own developers were concerned about oracles attitude that they went out to get third parties to add core components so that it could not be oracled :\
- atoponce 13y ago* ZFS is a volume manager. * ZFS is a RAID manager. * ZFS is also a filesystem. * Writes are handled in transaction groups (TXGs). * Every transaction group is written atomically. * ZFS keeps a revision history of the past 128 transactions written to disk. * ZFS is a Copy on Write filesystem. * As such, due to the previous 2 features, snapshots are free. * Snapshots are first class, read-only filesystems. * Snapshots can be upgraded to read-write clones. * Snapshots can be sent and received to other locations. * ZFS uses block-level deduplication. * ZFS supports transparent compression. * Every metadata and block data is checksummed with SHA256 by default. * Other checksum algorithms are supported. * ZFS uses a "slab allocator" to minimize fragmentation. * ZFS implements an "intent log" for synchronous writes. * The intent log can be migrated to a fast SSD or NVRAM drive. * ZFS uses advanced caching implementing for MRU/LRU and MFU/LFU caches. * A secondary cache (outside of RAM) can be installed on fast SSDs. * ZFS uses dynamic striping with its RAID arrays. * ZFS supports triple parity RAID. * ZFS autoheals bit rot when a block does not match its checksum, if the pool is redundant. * ZFS fully supports advanced format disks (4k blocks and beyond). * In fact, block sizes are dynamic from 512 bytes to 128K (or 1M in the proprietary ZFS). * In the proprietary release of ZFS, native encryption is supported. * In the Free Software release of ZFS, "feature flags" have been introduced to add on "plugins" without changing the core of the filesystem. * ZFS supports native NFS, allowing the mount to be available before the export. * ZFS supports native SMB for the same reason. * ZFS supports native iSCSI, also for the same reason. * ZFS can create static sized block devices called "ZVOLS". * ZFS pools can be exported and imported. * ZFS "scrubs" data to find blocks that do not match their checksum. * The Free Software release of ZFS is supported on GNU/Linux, OpenIndiana, SmartOS, FreeBSD, and many other operating systems. * Administration of ZFS is done via 3 commands: zpool(8), zfs(8) and zdb(8).
- midas007 13y agoGreat list. L2ARC, zil can each have their own volume configuration (mirror, etc.) For example, using different types of SSDs for each. http://forums.freenas.org/index.php?threads/zfs-and-ssd-cache-size-log-zil-and-l2arc.6345/ http://forums.freenas.org/index.php?threads/zfs-and-ssd-cach... zfs send & receive ... Send snapshots around like a fancy SAN. raidz (N+1 - like raid5) raidz2 (N+2 - like raid6) raidz3 (N+3) It's also way faster and cheaper to put together boxes from commodity enterprise server hardware, making hardware raid cards basically expensive shelf dust catchers along with overpriced SANs and NASes. (Extra shout out for iXsystems, not because they use lots of Python, but because of massive awesomeness supporting FreeBSD and FreeNAS. Also their parties put Defcon afterparties to shame.) Conclusion: Full ZFS is often better than a SAN, NAS and/or hardware solutions. Also protip: Direct attached is way, way faster than 10 GbE, FC or IB, especially if images are directly available to compute nodes.