16 ms·
Five Years of Btrfs
- e40 7y agoI've heard a lot of people say they won't use Btrfs due to reliability. Would have been nice to see that addressed.
- ailideex 7y agoWhat about the reliability? Are many people losing data with Btrfs?
- mdip 7y agoCaveat: I don't use RAID[0]. In about 4 years of running it on a couple of servers and countless virtuals/desktops, I've never had a reliability issue that was directly related to btrfs. I do not have my servers plugged in to UPSes, I have the occasional "shutdown due to power loss". The only time I've lost data has been due to cable disconnection in my hardware RAID array, and even then I was able to recover a substantial amount of its `btrfs` stored files. [0] Well, not filesystem-provided RAID; I have LSI controllers that provide the array to the OS as a single disk.
- londons_explore 7y agoMy understanding is most bugs are ironed out of btrfs itself, but tooling is still weak. For example, if you have a disk drive go bad on you and you manage to recover ~ half of the sectors with a disk imaging tool, you won't be able to extract files from the image without extreme effort.
- zerkten 7y agoWhy hasn't this caught up? Is it the case that data recovery companies are hoarding this after investing in their own tools, or something fundamental to the community?
- jbotz 7y agoAs best I can tell, reports of data loss on btrfs are all from the early 20-teens; after about 2014 or so I can't find anyone who claims to have lost data due to a btrfs bug on an up-to-date system.
- Fnoord 7y agoRAID5 on btrfs has a write hole last time I checked. Bug has been around forever, and was around in 2014 for sure. Phoronix has some thorough performance comparisons between Ext4fs, Btrfs, XFS, and ZFS.
- vetinari 7y agoThat one is here to stay, it is a property of software-based RAID. If it bothers you, use UPS.
- zielmicha 7y agoZFS-based RAID-5 (called raidz1) doesn't have write hole. https://blogs.oracle.com/ahl/what-is-raid-z https://blogs.oracle.com/ahl/what-is-raid-z
- vetinari 7y agoBecause ZFS raidz1 is not raid5, it's even labelled differently. Yes, it is a parity-based raid, but has slightly different semantics.
- Fnoord 7y agoHardware RAID can also suffer from this indeed but does ZFS suffer from it as well? With exactly the same impact? AFAIK the filesystem stays consistent on ZFS.
- vetinari 7y agoraidz1 is not raid5. From https://pthree.org/2012/12/05/zfs-administration-part-ii-raidz https://pthree.org/2012/12/05/zfs-administration-part-ii-rai... : > ather than the stripe width be statically set at creation, the stripe width is dynamic. Every block transactionally flushed to disk is its own stripe width. Every RAIDZ write is a full stripe write. Further, the parity bit is flushed with the stripe simultaneously, completely eliminating the RAID-5 write hole. So, in the event of a power failure, you either have the latest flush of data, or you don't. But, your disks will not be inconsistent. > There's a catch however. With standardized parity-based RAID, the logic is as simple as "every disk XORs to zero". With dynamic variable stripe width, such as RAIDZ, this doesn't work. Instead, we must pull up the ZFS metadata to determine RAIDZ geometry on every read. If you're paying attention, you'll notice the impossibility of such if the filesystem and the RAID are separate products; your RAID card knows nothing of your filesystem, and vice-versa. This is what makes ZFS win.
- zerkten 7y agoIt's the default for new Synology devices, and has been for a while. I suspect others are using it in a similar situation for home-grade NAS and up into the prosumer end of the market. I feel like Btrfs is probably going to be well tested here, but I wonder how many of these users are diagnosing Btrfs problems when they occur? It's going to be more evident to some people, and you have to assume that some of the vendors are competent, but this is against a backdrop of people throwing this kit away or starting from scratch versus performing a root cause analysis. I've personally been running this since it was stable on my DS1515+. I haven't had filesystem issues yet, but I make sure my important stuff is backed up elsewhere. A local backup like this is convenient for faster recovery in a lot of situations though which is why I keep it. I've SSH'd to the device and played around a little, but I fear I'd hit something proprietary, if the worst recovery situation occurred and I had to get everything from the DS1515+. If it was just an Ubuntu box I wouldn't have those fears, but the Syno NAS package is compelling.
- kasabali 7y agoReliability does not only mean data loss. It may not be losing data but crashing every few hours, or locking up the system, or requiring constant monitoring and maintenance etc.
- pojntfx 7y agoLove using Btrfs; the is no better filesystem than it nowadays that it's reliability issues have been fixed.
- mdip 7y agoI tend to agree with you here -- reliability has been a non-issue for me, though I've never configured `btrfs` in its RAID configuration. Performance becomes an issue in certain cases, but in every one that I've encountered, adjusting configuration has resolved the problems to my satisfaction. Would my Windows 10 VM run better under a different filesystem, rather than `btrfs` with various tweaks applied? Reading relatively recent articles on the subject would suggest that it would, however, I'd rather work with a single filesystem type and understand its strengths/weaknesses than manage two different filesystems as long as I can get performance to a usable state.
- ansible 7y agoWe have run btrfs in RAID configuration, but that has usability issues, even just doing RAID-1. We've switched back to using MD (mdadm) for RAID-1 setup, and then using btrfs on top of that for the snapshots, send / receive, block-level CRC and such. Dealing with failed drives isn't as easy with btrfs as it is with Linux MD.
- pmarreck 7y agoI have to beg to differ here as I had a different experience that I literally just posted about to Reddit yesterday https://www.reddit.com/r/zfs/comments/eu1qsj/a_tale_of_two_filesystems/ https://www.reddit.com/r/zfs/comments/eu1qsj/a_tale_of_two_f... tl;dr Unbeknownst to me I had a bad drive cable for an external NVMe enclosure that was causing intermittent I/O errors (only during high drive utilization) that went undetected by BTRFS and slowly corrupted my drive, eventually leading to an unbootable and unrepairable system (and to be fair, I should have scrubbed instead of attempting btrfsck --repair from another booted drive, but I don't care what you say, a --repair function should NOT potentially cause FURTHER corruption if it is at all available in the tooling! Like, just fucking rip it out if it can potentially make things worse, or recode the damn thing to just act defensively... jeez) Wiped the drive and started over with Ubuntu 19.10 and its new integrated ZFS on Root support... ZFS detected the IO issue pretty much instantly and prevented further errors by freezing I/O. Swapped the cable out during my troubleshooting and the issue went away. Also, drive is plenty fast, read test at 800MB/s
- curt15 7y agoBTRFS is well known for being ill-suited to VMs or databases. How come ZFS doesn't have that reputation?
- ailideex 7y agoZFS on Solaris 10 was ill suited for almost everything out of the box and was not feasible for MongoDB due to the half are integration between ZFS and the rest of Solaris
- xorcist 7y agoThere is an attribute called NOCOW that can be set on specific files that should not be copy-on-write, which is what messes with databases, filesystem images and other things that needs fast in-place updates. It can also be a set as a flag on subvolumes.
- jolmg 7y agoYou can also set such an attribute on files and subvolumes in btrfs. https://wiki.archlinux.org/index.php/Btrfs#Disabling_CoW https://wiki.archlinux.org/index.php/Btrfs#Disabling_CoW
- rayvd 7y agoI think ZFS was just ahead. We considered btrfs many years ago to serve up VM's for ESX via NFS, but it just wasn't as performant unless you ran in async mode which jeopardized data integrity. ZFS let you introduce SSD-based ZIL and L2ARC caching which made performance totally fine in sync mode. We're mostly NetApp AFF these days, but early on had close to a petabyte of ZFS-based storage power VM's on SuperMicro or Dell gear. Definitely was higher touch than NetApp but far less expensive.
- epx 7y agoI have been using btrfs in my "NAS"/personal server for 3 years, changed disk configuration a couple times, I do snapshots every hour and prune them using a Fibonacci-like timeline, no problems yet.
- Teknoman117 7y agoMy experience has been the same. Admittedly, I've not tried native BTRFS parity raid (I'm sitting the volume on top of mdraid). But, I ran the "mkfs.btrfs" 5 years ago at this point for my desktop and no data loss yet. I back things up religiously, so I'm not too worried about the volume failing, but it'll be nice if btrfs parity raid gets stabilized, because I could replace my current NAS storage config. I used to use ZFS on my NAS, but after running it for a year and fiddling with it, I wasn't able to tune it in a way I liked. I always had random performance problems and zvols were super slow. It's now dm-integrity on all disks, an mdraid raid6 volume over those, with LVM2 on top of that and mirrored NVMe disks as a read and write cache. I also wish BTRFS would add extents at some point so you could run virtual machine images from it without weird performance issues from time to time (although I imagine this is less of an issue on SSDs because they're "fragmented" inside anyways).
- lousken 7y agoDid anyone had the courage to use btrfs in production? Any stories to share?
- jhalstead 7y agoSeems like Facebook uses it: "Btrfs has played a role in increasing efficiency and resource utilization in Facebook’s data centers in a number of different applications. Recently, Btrfs helped eliminate priority inversions caused by the journaling behavior of the previous filesystem, when used for I/O control with cgroup2 (described below). Btrfs is the only filesystem implementation that currently works with resource isolation, and it’s now deployed on millions of servers, driving significant efficiency gains." https://engineering.fb.com/open-source/linux/ https://engineering.fb.com/open-source/linux/
- alexgartrell 7y agoYeah there are a remarkable set of container runtime tasks (package downloads, rootfs creation and management, etc) that are way easier with btrfs. It wasn’t always smooth sailing but luckily Chris, Josef, Omar and others are awesome and now (and for the last while) we are asking for features rather than fixes.
- agravier 7y agoIt's the default on recent Synology NAS, in my experience. No particular issue in my limited experience. Mostly transparent for the user.
- barclay 7y ago(also a happy syno user here, been using it on several NAS's quite happily). My rough understanding is synology did some pretty heavy modifications to btrfs in their implementation though... (a quick google finds me nothing to back this up, but i remember reading about it somewhere...)
- InTheArena 7y ago
- mdip 7y agoI've been a `btrfs` user for the better part of 4 years despite, at the time, a very vocal group providing advice against it[0]. I'll be the first to say that it isn't a silver bullet for everything. But then, what filesystem really is? Filesystems are such a critical part of a running OS that we expect perfection for every use case; filesystem bugs or quirks[1] result in data loss which is usually Really Bad(tm). That said, for the last two years, I've been running Linux on a Thinkpad with a Windows 10 VM in KVM/qemu -- both are running all the time. When I first configured my Windows 10 VM, performance was brutal; there were times when writes would stall the mouse cursor and the issue was directly related to `btrfs`. I didn't ditch the file-system, I switched to a raw volume for my VM and adjusted some settings that affected how `btrfs` interacted with it. I discovered similar things happened when running a `balance` on the filesystem and after a bit of research, found that changing the IO scheduler to one more commonly used on spindle HDDs made everything more stable. So why use something that requires so much grief to get it working? Because those settings changes are a minor inconvenience compared against the things "I don't have to mess with" to cover a bigger problem that I frequently encountered: OS recovery. An out-of-the-box OpenSUSE Tumbleweed installation uses `btrfs` on root. Every time software is added/modified, or `yast` (the user-friendly administrative tool) is run, a snapshot is taken automatically. When I or my OS screws something up, I have a boot menu that lets me "go back" to prior to the modification. It Just Works(tm). In the last two years, I've had around 4-5 cases where my OS was wrecked by keeping things up to date, or tweaking configuration. In the past, I'd be re-installing. Now, I reboot after applying updates and if things are messed up, I reboot again, restore from a read-only snapshot and I'm back. I have no use for RAID or much else[2] which is one of the oft-repeated "issues" people identify with `btrfs`. It fits for my use-case, along with many of the other use-cases I encounter frequently. It's not perfect, but neither is any filesystem. I won't even argue that other people with the same use case will come to the same conclusion. But as far as I'm concerned, damn it works well. [0] I want to say that an installation of openSUSE ended up causing me to switch to `btrfs`, but I can't remember for sure -- that's all I run, personally, and it is a default for a new installation's root drive. [1] Bug: a specific feature (i.e. RAID) just doesn't work. Quirk: the filesystem has multiple concepts of "free space" that don't necessarily line up with what running applications understand. [2] My servers all have LSI or other hardware RAID controllers and present the array as a single disk to the OS; I'm not relying on my filesystem to manage that. My laptop has a single SSD.
- zozbot234 7y agoI'm so sorry teacher, Btrfs ate my homework.
- alyandon 7y agoI use btrfs in raid1 mode and the ability to shrink/grow/add/remove devices at will without data loss or extended downtime led me to choose btrfs over zfs on my home servers.
- cyphar 7y agoYou can grow and add/remove raid1 devices (mirror vdevs) in ZFS without any significant work or downtime. Shrinking does require a bit more work, but depending on your setup it can be done fairly painlessly with send/recv (and shrinking is usually not something which is a very common administrative operation).
- 3fe9a03ccd14ca5 7y agoHow? My understanding is that you create a new vdev and add the old vdev as a device, basically recursively creating volumes with each new device you add.
- deleted 7y ago[deleted]
- cyphar 7y agoWhich operation are you asking about? [1] is a sister comment which I posted that outlines how to do most of the operations I mentioned. [1]: https://news.ycombinator.com/item?id=22168494 https://news.ycombinator.com/item?id=22168494
- xzcat 7y ago"fairly painlessly" and "without significant work or downtime" doesn't sound like it lines up with btrfs's, which I would describe as "one command and zero downtime (just some io load if you rebalance immediately)" for both operations. btrfs is also mainline, which increases how painless it is to use. BTRFS does have some scary stories from earlier in its development, and true raid5 seems like it's unlikely to be safe for quite a while, but raid1 and "normal" fs usage has been rock solid in my experience. The only time I've ever had an issue was probably 4 years ago at this point, and it was solved by just booting an Arch live iso and running a btrfs command that was basically "fix exactly the bug that your error message indicates". I don't remember exactly what it is, something about two sizes not matching, but googling the text it showed at boot led me directly to the command to fix it. Certainly dramatically less trouble than I've ever had when hardware RAID goes south. I do agree that modern lvm does probably compete with btrfs, but again you're trading how dang simple btrfs raid1 is to manage for monkeying with partitions in lvm in exchange for ~some? performance. IMO ZFS is in a weird spot where I don't know where I'd use it. It's too complicated/annoying to admin for me to want to run it in my basement for myself/my family, and for anything bigger or more professional I'd use ceph or a problem-domain-specific storage system (HDFS, clickhouse, aws, etc).
- nickik 7y agoBeing 'The Dude' of file system is literally the opposite of what I want. When looking at ZFS talks and the incredible complexity of some of those operations that Btrfs seems to think are 'no big deal', I will simply not trust that. Specially because it has been proven over and over again that Btrfs claims its 'stable' and then a new series of issues show up. Or its 'stable' but not if you use 'XY feature', or if the disk is 'to full' or whatever. I remember using it after I had heard it was 'stable' and it eat my data not long after (not using crazy features or anything). I certainty will not use it again. A FS should be stable from the beginning, as stable core that you can then build features around, rather then a system with lots of feature that promises to be stable in a couple years (and then wasn't years after being in the kernel already). Using ZFS for me has been nothing but joy in comparison. Growing the ZFS pool for me has been no issue at all, I never saw a reason why I would want to reconfigure my pool. I went from 4TB to 16TB+ so far in multiple iterations. Overall not having ZFS in Linux is a huge failure of the Linux world. I think its much more NIMBY then a license issue.
- _jal 7y agoI mostly agree, and that's largely why I use ZFS a lot. But: > A FS should be stable from the beginning, If this is your standard, I don't think there's a file system out there that meets it. ZFS has had data-loss bugs. I doubt there is any non-toy file system that hasn't. I've thought about what standard should apply to this - it is a prove-a-negative problem, that filesystem-X in combination with whatever recent kernel will not lose data. I don't have a good answer, but the one I came up with is "multiple years without a dataloss bug, of quick turnaround to other bug fixes, and a warm-fuzzy feeling about the developers."
- nickik 7y agoIt was designed to primary not lose data from the very beginning. That was at the very core of every design choice. Maybe there were a few such bugs but I have not read of any, while in comparison Btrfs I have read a whole of them. Compare how bcachefs/zfs approaches these challenges and then go back to the early years of Btrfs. There is really no comparison.
- Shalle135 7y agoIs there any specific reasons to run btrfs over for example ext4? You can create/shrink/grow pools, create encrypted volumes etc by using LVM. It all depends on the application but in the majority of cases the io performance of btrfs is worse than the alternatives. Redhat for example choose to deprecate btrfs for unknown reasons while SUSE made it it’s default. The future of it seems uncertain which may cause a lot of headache’s in major environments if implemented there.
- derefr 7y agoRedhat and SUSE (SLES) are both enterprise environments, so at every level, they have to choose one tech stack to go all-in on (i.e. to train their support staffs on), and then discourage their customers from using the others. (“Deprecating” a component, for such orgs, means that some of their customers are now stuck with it, and they’ll continue to support those customers in their use of it, but they certainly won’t support new customers using it.) The fact that one enterprise-support provider went all-in on Btrfs, while another didn’t, basically tells you that the choice is pretty arbitrary. If no enterprise-support provider used Btrfs, then I’d be concerned.
- Arnavion 7y agoThe enterprise provider that actually develops btrfs continues to support btrfs, and one enterprise provider that doesn't stopped supporting it. People treat RH stopping support of btrfs as some sort of death knell for it. Meanwhile all the btrfs users are confused why RH's opinion should matter at all when they weren't that involved with developing it in the first place. As an opensuse user, btrfs has saved multiple machines from botched updates by letting me revert to the snapshot from right before the update was applied (opensuse's update tool automatically takes snapshots before and after updates).
- Conan_Kudo 7y agoRed Hat used to be heavily involved in Btrfs development. In fact, they are present in a huge chunk of its development in the first few years. But their developers were hired away by Facebook, leaving Red Hat with nobody who work on Btrfs regularly. That's the underlying cause for why they stopped supporting it. Hiring someone to work on Btrfs takes time and effort that they don't have a reason to spend right now.
- derefr 7y agoA question for HN: what filesystem and/or block-device abstraction layer would you use on a database server, if you wanted to perform scheduled incremental backups using filesystem-level consistent snapshotting and differential snapshot shipping to object storage, instead of using the DBMS’s own replication layer to achieve this effect? (I.e. you want disaster recovery, not high availability.) Or, to put that another way: what are AWS and GCP using in their SANs (EBS; GCE PD) that allows them to take on-demand incremental snapshots of SAN volumes, and then ship those snapshots away from the origin node into safer out-of-cluster replicated storage (e.g. object storage)? It it proprietary, or is it just several FOSS technologies glued together? My naive guess would be that the cloud hosts are either using ZFS volumes, or LVM LVs (which do have incremental snapshot capability, if the disk is created in a thin pool) under iSCSI. (Or they’re relying on whatever point-solution VMware et al sold them.) If you control the filesystem layer (i.e. you don’t need to be filesystem-agnostic), would Btrfs snapshots be better for this same use-case?
- StreamBright 7y ago>> Or, to put that another way: what are AWS and GCP using in their SANs (EBS; GCE PD) that allows them to take on-demand incremental snapshots of SAN volumes, and then ship those snapshots away from the origin node into safer out-of-cluster replicated storage (e.g. object storage)? As far as I know AWS does not use SANs because they consider it as anti-pattern. Most backups land on S3 because of reliability and price.
- polskibus 7y agoSo how is S3 implemented? Does it reuse any publicly available open source component?
- derefr 7y agoI don’t think they’ve published anything specifically on S3’s architecture (someone please correct me if I’m wrong, I last looked into this a long time ago), but 1. they came out with S3 soon after coming out with their Dynamo paper (before releasing DynamoDB, even); and 2. there’s a good constructive proof, as a studyable FOSS system, for how to build object storage on top of a Dynamo architecture, in the form of Riak CS (object storage) which is built atop Riak KV (a Dynamo impl.) Riak CS seems to make pretty much the same set of guarantees (in terms of time/space complexity of operations, possible durability numbers per scaled number of copies, etc.) that S3 does, so it’s a fair guess that they’re similarly-architected systems.
- InTheArena 7y agoI went on a quest a few years ago, thinking it would be good for the industry to standardize on a single next generation filesystem for UNIX. I started with ZFS on linux since that seemed to have the most vocal advocates. That lasted about a half year, until a bug in the code resulted in a completely corrupt disk, and I had to restore 4TB of data over a month from offside backups. That plus the licensing confusion around ZFS has made it impossible for ZFS to be the defacto choice. I went down the BTRFS path, despite it's dodgy reputation when netgear announced their little embedded NASes, and switched my server over to it. The experience was solid enough that I bought high-end synology and have had zero problems with it.
- clSTophEjUdRanu 7y agoI really don't understand the insane hype around ZFS. You can't read any thread that touches on filesystems without the ZFS zealots coming out.
- asveikau 7y agoI don't think I am a zealot, nor a heavy user, but I use it on 1 machine at home (an NFS server running FreeBSD, which I have clients for elsewhere in my house). I came to this idea when I saw some data loss on some magnetic disks in my house, and repairing or even assessing the level of damage was difficult. My experience is that it's pretty good. The tooling does what it says without a lot of drama. I can scrub while the system is in use and don't notice it mostly. I have seen some small corruptions that it was able to flag for me with specific filenames and fix. Snapshotting and send/receive is also very handy. I heard some people say they don't like to use it under heavy load. That seems reasonable to me. You're paying costs to get the integrity piece. So it's not for every use or every user. It is very good at what it does, however.
- stiray 7y agoSame with me. I just figured out at some point, 10 years ago, that it is nice to have snapshots on root disk. And figured out FreeBSD is supporting ZFS. Tryed it, loved it, used it. The ZFS on linux was destabilized in latest versions (`ls /.zfs/snapshots`) and they blew it considerably by adding it to systemd (I need to reboot fedora multiple times before it boots ever since), but at least I know that my data are not lost (unlike btrfs, had two major crashes in two years). Quite frankly I'll rather wait for Raisser to get out of jail than use btrfs again. Anyway, I bet on Hammer2.
- geophertz 7y agoIs using btrfs on a personal machine something to do? It seems that all the comments as well as articles about it, just assume you're running it on a server. The ability to add and remove disks on a desktop machine is very tempting.
- wtfrmyinitials 7y agoI've been running it on my desktop for a while and it's been wonderful. I have a cron job set to take a snapshot of the filesystem hourly so if I ever blow a file away or a package upgrade goes wonky I'm back up and running in minutes.
- izacus 7y agoIt's also worth noting that Synology uses btrfs as an option to do checksumming and snapshots on their NAS devices. They're still using their own RAID layer though.
- ValentineC 7y ago> They're still using their own RAID layer though. Synology's RAID implementation is largely mdadm + LVM.
- abotsis 7y agoIt’s worth noting that much of the premise of the article (wanting flexibility) is outdated. Zfs has support for removing top-level raid 0/1 vdevs now. So you can take a raid10 pool, and remove a top level mirror vdev completely. Note that this doesn’t work for raid5/6 vdevs, but as the author points out, those are becoming less and less used because of rebuild time and performance. In addition to the slew of other features Btrfs is missing (send/recv, dedup, etc) zfs allows you to dedicate something like an Intel optane (or other similar high write endurance, low latency ssd) to act as stable storage for sync writes, and a different device (typically mlc or tlc flash) to extend the read cache.
- abotsis 7y ago*Just kidding on send/recv, looks like it’s there now. Substitute with encryption if you need another example.
- herf 7y agozfs remove is not a very good implementation - it keeps the old blocks around (as a virtual device) and redirects them to new locations. This is fine for "oops I accidentally added a device" but not great otherwise.
- kstrauser 7y agoI think there's a selection bias here: people using RAID 5/6 may not be using ZFS as much because it's not well supported. I'd bet money that those levels are much more common in SOHO settings than RAID 10 is, because it's still the sweet spot for "I need lots of storage" vs "...and am willing to spend drive's worth of storage on availability". For instance, anyone using a NAS primarily as a backup target for desktops and small servers may love RAID 5, but be unwilling to throw money at a "better" RAID 10 setup.
- the8472 7y agobtrfs has send/recv. And dedup, which is more efficient that ZFS' since it can be performed offline, on select parts of the filesystem and doesn't have to keep gigabytes of dedup tables in memory.
- lazylizard 7y ago¯\_(ツ)_/¯ Raidz2+spares, compression, snapshots and send/receive are very useful. And zil and cache are easier than lvmcache..
- zielmicha 7y agofsync is still a bit slow on BTRFS (on ZFS too, but to a smaller degree). For example, I just did a quick benchmark on Linux 5.3.0 - installing Emacs on fresh Ubuntu 18.04 chroot (dpkg calls fsync after every installed package). ext4 - 33s, ZFS - 50s, btfrs - 74s (test was ran on Vultr.com 2GB virtual machine, backing disk was allocated using "fallocate --length 10G" on ext4 filesystem, the results are very consistent)
- c0ffe 7y agoI have a small Nextcloud instance at home that uses BTRFS (on HDD, with noatime option) for file storage, and XFS (on SSD) for database. I started it just for testing, and has been running for up to two years, and had no problems so far.
- tezzer 7y agoI've had one issue with btrfs that took it off my radar completely. A customer had a runaway issue that filled a btrfs device with unimportant things. We found the errant process and killed it, but apparently if a btrfs device is completely full, you can't delete anything to free up space. File removal requires some amount of free space. Bricked the device, annoyed a customer, back to ext4.
- takeda 7y agoZFS had this issue (I believe fixed) workaround was to pick up one large file that you wanted to delete and do `echo -n > /the/unimportant/file` once the file was reduced in size to 0, rm started to work again. Not sure if that workaround would work in btrfs, but it worked on ZFS.
- rcthompson 7y agoWhat happens if the file has already found its way into a snapshot? Then presumably that command will not free any space.
- loeg 7y agoSee also: 'truncate -s 0 /the/file'
- pfranz 7y agoYep. I had this happen a few weeks ago (I'm not sure how much maintenance the server has had since it was set up 2-3 years previous). Thankfully, after seeing whatever the error was ("No space left on device" or something) and furrowing my brow it seemed obvious enough to try without having to search for a solution. It seemed just dumb enough to work.
- gravypod 7y agoI've seen a lot of the hacker community focusing on btrfs and zfs but very little focusing on ceph. I think ceph has a lot of the features that we want in a file system and some things that aren't even possible on traditional file systems (per-file redundancy settings) with very little downsides. The setup is a little more complex involving a few daemons to manage disks, balance, monitor, etc. I wish there was something similar to FreeNAS for ceph that only focused on making the experience seemless because I think if it became more popular in the home lab space we'd see lots of cool tools pop up for it.
- louwrentius 7y agoI love Ceph, I even wrote an intro about it for those who are not familiar with it. https://louwrentius.com/understanding-ceph-open-source-scalable-storage.html https://louwrentius.com/understanding-ceph-open-source-scala... But Ceph is not designed to be a competitor to BTRFS or ZFS. The core vision of Ceph is scalability. If you need petabytes of storage and the performance to scale with it, take a look at Ceph. I may be totally wrong here, but from what I understand about Ceph, it's not meant as a file system for a single computer. I don't understand the idea of running Ceph on your laptop/desktop. It's possible to run it that way but it defeats it's purpose. I've build a small lab setup with Ceph: https://louwrentius.com/my-ceph-test-cluster-based-on-raspberry-pis-and-hp-microservers.html https://louwrentius.com/my-ceph-test-cluster-based-on-raspbe... Also, there's the issue of performance, in particular latency. That's a bit of a weak spot of Ceph, from what I can tell. Again, may be wrong. But I found these notes interesting. https://yourcmc.ru/wiki/Ceph_performance https://yourcmc.ru/wiki/Ceph_performance
- seabrookmx 7y agoThis. In fact, it's really common to use a ZFS array on single nodes, and then create a SAN using multiple such machines by layering Ceph on top.
- louwrentius 7y agoThat's interesting, but it's layers upon layers... (RIP latency), I think. Unless it's about just bandwidth and volume, then latency is not that big of a deal.
- kiney 7y agoI use BTRFS on several devices for years. The tooling is a bit rough, but no major problems. Just recently data checksumming saved me: In December I replace an old 2TB drive in my RAID1 (2+4+4+4) with an 8TB drive. The new drive had checksum errors after a few weeks which BTRFS handled gracefully. With "classical" RAID i might only have noticed when it's to late. (I RMAed the bad drive) [/dev/mapper/h4_crypt].write_io_errs 0 [/dev/mapper/h4_crypt].read_io_errs 0 [/dev/mapper/h4_crypt].flush_io_errs 0 [/dev/mapper/h4_crypt].corruption_errs 0 [/dev/mapper/h4_crypt].generation_errs 0 [/dev/mapper/h2_crypt].write_io_errs 0 [/dev/mapper/h2_crypt].read_io_errs 30 [/dev/mapper/h2_crypt].flush_io_errs 0 [/dev/mapper/h2_crypt].corruption_errs 0 [/dev/mapper/h2_crypt].generation_errs 0 [/dev/mapper/h1_crypt].write_io_errs 0 [/dev/mapper/h1_crypt].read_io_errs 0 [/dev/mapper/h1_crypt].flush_io_errs 0 [/dev/mapper/h1_crypt].corruption_errs 0 [/dev/mapper/h1_crypt].generation_errs 0 [/dev/mapper/h3_crypt].write_io_errs 0 [/dev/mapper/h3_crypt].read_io_errs 0 [/dev/mapper/h3_crypt].flush_io_errs 0 [/dev/mapper/h3_crypt].corruption_errs 0 [/dev/mapper/h3_crypt].generation_errs 0 [/dev/mapper/luks-e120f41e-9c8a-4808-876f-fa6665ee8bb8].write_io_errs 0 [/dev/mapper/luks-e120f41e-9c8a-4808-876f-fa6665ee8bb8].read_io_errs 16 [/dev/mapper/luks-e120f41e-9c8a-4808-876f-fa6665ee8bb8].flush_io_errs 0 [/dev/mapper/luks-e120f41e-9c8a-4808-876f-fa6665ee8bb8].corruption_errs 20619 [/dev/mapper/luks-e120f41e-9c8a-4808-876f-fa6665ee8bb8].generation_errs 0 edit: formatting
- gitgudnubs 7y agoStorage spaces is probably the best software raid available today. Unfortunately, it comes with windows. It supports heterogenous drives, safe rebalancing (create a third copy, THEN delete the old copy), fault domains (3-way mirror, but no 2 copies can be on the same disk/enclosure/server/whatever), erasure coding, hierarchical storage based on disk type (e.g., use NVMe for the log, SSD for the cache), clustering (paxos, probably). Then you toss ReFS on top, and you're done. The only compelling reasons to buy windows server are to run third party software or a storage spaces/ReFS file share.
- williesleg 7y agoButters FS! Yay!
- shmerl 7y agoI'm using Btrfs currently, but I'm waiting for Bcachefs to replace it.
- mekster 7y agoHow far has it come to replace any of the production ready filesystems? It says it's feature complete in 2015 and was trying to put itself into kernel mainline in 2018 but I don't see much about anyone using it in production.
- shmerl 7y agoIt's not production ready so far for sure. I didn't really follow all details on that.
- cmurf 7y agokernel 5.5 released Sunday. Btrfs now has raid1c3, raid1c4 profiles for 3 and 4 copy raid1. Adds new checksum algorithms: xxhash, blake2b, sha256. Async discards coming in 5.6. https://lore.kernel.org/linux-btrfs/cover.1580142284.git.dsterba@suse.com/T/#u https://lore.kernel.org/linux-btrfs/cover.1580142284.git.dst...
- pQd 7y agoi've been using BTRFS since 2014 to store backups. there is a noticeable performance penalty when rsync'ing hundreds of thousands of files to a spinning-rust disk connected to USB-SATA dock when BTRFS is used instead of EXT4. i'm accepting it in exchange for ability to run scheduled scrub of the data to detect potential bitrot. since 2017 i'm also using BTRFS to host mysql replication slaves. every 15 min, 1h, 12h crash-consistent snapshots of the running database files are taken and kept for couple of days. there's consensus that - due to its COW nature - BTRFS is not well suited for hosting vms, databases or any other type of files that change frequently. performance is significantly worse compared to EXT4 - this can lead to slave lag. but slave-lag can be mitigated by: using NVMe drives and relaxing durability of MySQL innodb engine. i've used those snapshots few times each year - it worked fine so far. snapshots should never be the main backup strategy, independently of them there's a full database backup done daily from masters using mysqldump. snapshots are useful whenever you need to very quickly access state of the production data from few minutes or hours ago - for instance after fat fingering some live data. during those years i've seen kernel crashes most likely due to BTRFS but i did not lose data as long as the underlying drives were healthy.
- cyphar 7y agoThis article makes a few mistakes with regards to ZFS. Some are understandable (the author presumably last looked at the state of ZFS 5 years ago), but some were not true even 5 years ago: > If you want to grow the pool, you basically have two recommended options: add a new identical vdev, or replace both devices in the existing vdev with higher capacity devices. You can add vdevs to a pool which are different types or have different parities. It's not really recommended because it means that you're making it harder to know how many failures your pool can survive, but it's definitely something you can do -- and it's just as easy as adding any other vdev to your pool: % zpool add <pool> <vdev> <devices...> This has always been possible with ZFS, as far as I'm aware. > So let’s say you had no writes for a month and continual reads. Those two new disks would go 100% unused. Only when you started writing data would they start to see utilization This part is accurate... > and only for the newly written files. ... but this part is not. Modifying an existing file will almost certainly result in data being copied to the newer vdev -- because ZFS will send more writes to drives that are less utilised (and if most of the data is on the older vdevs, then most reads are to the older vdevs, and thus the newer vdevs get more writes). > It’s likely that for the life of that pool, you’d always have a heavier load on your oldest vdevs. Not the end of the world, but it definitely kills some performance advantages of striping data. This is also half-true -- it's definitely not ideal that ZFS doesn't have a defrag feature, but the above-mentioned characteristic means that eventually your pool will not be so unbalanced. > Want to break a pool into smaller pools? Can’t do it. So let’s say you built your 2x8 + 2x8 pool. Then a few years from now 40 TB disks are available and you want to go back to a simple two disk mirror. There’s no way to shrink to just 2x40. This is now possible. ZoL 0.8 and later support top-level mirror vdev removal. > Got a 4-disk raidz2 pool and want to add a disk? Can’t do it. It is true that this is not possible at the moment, but in the interest of fairness I'd like to mention that it is currently being worked on[1]. > For most fundamental changes, the answer is simple: start over. To be fair, that’s not always a terrible idea, but it does require some maintenance down time. This is true, but I believe that the author makes it sound much harder than it actually is (it does have some maintenance downtime, but because you can snapshot the filesystem the downtime can be as little as a minute): # Assuming you've already created the new pool $new_pool. % zfs snapshot -r $old_pool/ROOT@base_snapshot % zfs send $old_pool/ROOT@base_snapshot | zfs recv $new_pool/ROOT # The base copy is done -- no downtime. Now we take some downtime by stopping all use of the pool. % take_offline $old_pool # or do whatever it takes for your particular system % zfs mount -o ro $old_pool/ROOT # optional % zfs snapshot -r $old_pool/ROOT@last_snapshot % zfs send -i @base_snapshot $old_pool/ROOT@last_snapshot | zfs recv $new_pool/ROOT # Finally, get rid of the old pool and add our new pool. % zpool export $old_pool % zpool import $new_pool $old_pool % zfs mount -a # probably optional [1]: https://www.youtube.com/watch?v=Njt82e_3qVo https://www.youtube.com/watch?v=Njt82e_3qVo