6 ms·
This is symptomatic with one of my two big problems with journalling oriented file-systems. My problems with journalling are two fold: 1) They are very slow:
by cokernel_hacker 14y ago
This is symptomatic with one of my two big problems with journalling oriented file-systems.
My problems with journalling are two fold:
1) They are very slow:
1a) You have a nice big sequential write into the journal, which is OK.
1b) A flush track cache to make sure it is actually in the journal. This can sync whatever has accumulated in the track cache which might not just be journal data.
1c) The actual overwrites that spew data randomly over the drive.
1d) Writes to update the journal header/terminate the transaction.
1e) A final flush track cache which will sync who-knows-what onto the platters/flash.
2) Replay behavior of the journal log is _very_ fragile code. You need to handle lots of terrible cases, the most awful of which is to ensure that you don't play older transaction on top of newer transactions. You might say "hey, that shouldn't happen!" but it happens. It happens because the code is not trivial to write and detecting these bad cases aren't trivial. Even if you do get the code write, you are still screwed. Why? Because if your drive does not support the flush track cache mechanism you are in for a world of pain. You can have a journal and journal header that is ancient if it just stuck in some cache...
The ext* family of filesystems do not appear to have natural resiliency to this sort of problem. Instead, it appears to be a coordinated, concerted effort between various parts of the journaling code.
- riledhel 14y agoWhat is your main use case? What is your choice instead of ext*?
- cokernel_hacker 14y agoI like write-anywhere/COW filesystems. Unfortunately the open source ones kinda blow due to interesting performance penalties (I'm looking at you, btrfs back ref management) or crazy memory usage/non-extent based systems/read-modify-write induced mania (I'm looking at you, ZFS block tree)
- mcpherrinm 14y agoWhat non-open source ones do you like, then?
- sliverstorm 14y agoOh, you've probably never heard of them.
- zanny 14y agoI've been playing with btrfs under the 3.7 prerelease kernel on an arch kvm with a virtio disks, and with lzo compression btrfs is kicking some serious read / write butt vs every other major filesystem I can think of. Given sufficient hardware, of course. Haven't tried it on an ssd though, only on 7200rpm turnstables.
- lmm 14y agoI still use reiserfs for all my linux machines (ext2 for /boot); ZFS on freebsd. I haven't experienced the corruption I've seen under ext3 with either of them.
- rbanffy 14y ago> The ext* family of filesystems do not appear to have natural resiliency to this sort of problem. I believe the fact data corruption on ext4 is rare shows it's not really a huge problem. The only time I lost data on an ext4 filesystem was when I mistyped a wildcard for rm.
- cokernel_hacker 14y agoI would rather not have byzantine relations between fragments of code make up the policy of my file system's metadata resiliency thank-you-very-much. I would rather prefer that stale journal replays be no-ops, even at the expense of making journal replay slower.
- rbanffy 14y agotune2fs -O ^has_journal /dev/sda1 You're welcome.
- gjm11 14y ago> This is symptomatic with one of my two big problems with journalling oriented file-systems. Personally, I think it's more symptomatic of the problems of filesystem-oriented journalists. See the comment quoted by sciurus. (Irrelevant note: I really don't recommend reversing all the arrows in a Unix kernel. But I like your username anyway.)
- cokernel_hacker 14y agoMy apologies but what sciurus might be addressing and what I am talking about are addressing two very different things. There is a bug in ext4, https://lkml.org/lkml/2012/10/25/521 https://lkml.org/lkml/2012/10/25/521 Note that this is dated after the "Wed, 24 Oct 2012 17:31:29 -0400" post and talks about changing code related to a truncated journal. My post does not say anything about the likelihood of the dataloss due to such bugs, I mention what happens in every design in a WAL journaling system. Talking about the impact of a particular ext bug is up to the people who post on the LKML. To be honest, I don't really care how severe the bug is, I care about why there are bugs and journal truncation is a source of them.
- tytso 14y ago>Because if your drive does not support the flush track cache mechanism you are in for a world of pain. If your drive does not support CACHE_FLUSH mechanism, then pretty much any file system that are supposed to handle unclean shutdowns (ZFS, btrfs), etc., are going to be screwed. More generally, any time you have a file system, there will always be very delicate code. Fundamentally file systems are _hard_. It has to handle a huge amount of concurrent operations, and users want speed, so we use fine-grained locking, and there's a reason why file systems take 100 man years or so to become production ready. The btrfs folks started in 2007, and I warned them it would probably be at least 5-7 years, minimum before it would be ready for the enterprise. And now here it is five years later, and the community distributions (not the enterprise distros) are just starting to adopt btrfs. ZFS took seven years to develop before Sun announced it, and then it took a few more years before people trusted it on production servers. No one believes how hard it is when they start, but the history is pretty clear.