6 ms·
NixOS on Btrfs+tmpfs
- yjftsjthsd-h 4y agoNeat:) I would never use btrfs myself[0], but very happy to see people exploring all variations of these ideas. The one thing that's starting to bug me though, as I read blog posts about installing nixos: why is the install process so imperative/non-declarative? Once the system is up, the whole thing fits in configuration.nix, but to get there we still have to use masses of shell commands. Is anyone working on bridging that last gap and supporting partitions, filesystems, and mounts (I think that's all that's left?) from nix itself? [0] I lost 2 root filesystems to btrfs, probably because it couldn't handle space exhaustion. I'm paranoid now.
- rrix2 4y agoI have a fork of `justdoit.nix' https://github.com/cleverca22/nix-tests/blob/master/kexec/justdoit.nix https://github.com/cleverca22/nix-tests/blob/master/kexec/ju... which generates installer images with an auto-partitioning script and stub "configuration.nix" embedded in it with basically just enough of a system to get nixops or morph to deploy to it. it's kind of a pain to get that working the first few times since you have to wait for an image to bake and then test it in QEMU, and then make changes for nvme, etc, but it's brought up three systems now.
- yjftsjthsd-h 4y agoThanks! I'll have to give it a spin:)
- nix23 4y agoI loose every two years (when i test it again) a volume to btrfs, last time 2 month ago, with that simple "trick": -Fill your rootpartionion as root with "dd if=/dev/urandom of=./blabla bs=3m" -rm blabla && sync (we don't want to be unfair to such a fragile system) -Reboot and end up with unbootable / It's a mess, for a filesystem i would declare it as alpha stage.
- londons_explore 4y agoAll these "clever" filesystems can never guarantee not to run out of space for their own metadata. That's because even to delete a file they might need more space in the journal, or to un-copy-on-write some metadata. The mistake however is that even though it isn't practical to make theoretical guarantees that the filesystem won't end up full and broken, it is very possible to make such a thing only happen in exceeding unlikely cases. One runaway dd isn't that...
- KingMachiavelli 4y agoWhy can't they? For example, Btrfs reserves some storage for it's internal use which should be more than enough to update the journal to fix a full filesystem.
- londons_explore 4y agoCalculating exactly how much you need to reserve for the worst case is a near-impossible task. For example, say you try to delete a file, which is part of one of multiple identical snapshots, so deleting the file doesn't free up any space, but does require extra metadata to be written (since a new directory entry will be needed that shows the file is deleted in this snapshot only). The same operation could be done for millions of files, eating up all the reserved space. End result: full disk and unusable filesystem, even for deletes. The alternative is not to allow file deletes to use reserved space. But now when you have a full disk, some things become 'undeletable', since the only way to free space is to delete all copies of the file, but it isn't permitted to delete any one copy of the file since the intermediate state would use more disk space.
- cmurf 4y agoWhat is supposed to happen is the metadata commit fails due to enospc before the super block update. Thus the current super points to the current value working tree roots, not the partial/failed tree roots. Btrfs won't issue the writes for super block update until the device says the current metadata transaction is successfully on stable media. It is possible the filesystem is completely consistent (can be mounted, btrfsck finds no error), and yet not bootable due to the interruption of updates. Software updates are one transaction in user space but not atomic unless expressly designed for it. From the fs point of view, a software update might be broken up into dozens of fs transactions. It's also possible the device lies about writes being on stable media. If the fs writes some metadata, does flush/FUA, then super block write, and flash/FUA, the device should only write the super block after the prior write is on stable media. If it says the first flush succeeded but that write is still happening, and the super block write goes to stable media before all the metadata writes get to stable media and there's a crash or power failure, then you can in fact have a broken filesystem. The super points to tree roots that don't exist. This is definitely a device flaw not an fs flaw. Btrfs super blocks contain 3 backup roots. So it's possible to revert to an older and hopefully correct metadata generation (seconds to a couple minutes ago). But this has limited recovery potential. It's also completely thwarted right now if you use any discard mount option on an SSD because discard will ask the device to garbage collect recently freed metadata blocks. So the backup root trees pointed to by the super may already be zeros when they're needed. But any need for backup roots already means some kind of device (firmware) flaw.
- viraptor 4y ago> Most subvolumes can be mounted with noatime, except for /home where I frequently need to sort files by modification time. That doesn't sound right. Noatime turns off recording of the last access time, not modification.
- pdenton 4y agoIIRC, noatime is useful everywhere except /var/spool or /var/mail where certain daemons may depend on access time correctness.
- londons_explore 4y agoYou can't depend on access time correctness... Because someone else can come along and grep through all your files and now they're all accessed right now.
- traceroute66 4y ago> Most subvolumes can be mounted with noatime This noatime thing is an old-wive's tale that needs to die. AFAIK, most "modern" filesystems (XFS,BTRFS etc.) all default to relatime relatime maintains atime but without the overhead EDIT TO ADD: Actually,I've just done a bit of searching .... relatime has been the kernel mount default since >= 2.6.30 ! [1] [1] https://kernelnewbies.org/Linux_2_6_30 https://kernelnewbies.org/Linux_2_6_30 (scroll to 1.11. Filesystems performance improvements)
- 4y ago
- genghizkhan 4y agoI would prefer to do this on zfs, for which there is a lovely installation guide on the openzfs docs site. https://openzfs.github.io/openzfs-docs/Getting%20Started/NixOS/Root%20on%20ZFS.html https://openzfs.github.io/openzfs-docs/Getting%20Started/Nix...
- kaba0 4y agoI can vouch for how good it works. Using it on my personal laptop.
- grumpyprole 4y agoI tried ZFS-on-linux with Ubuntu 21.10 and it ate my data (ZFS panics when accessing certain files). Sure, Ubuntu does have a habit of using unstable kernels, but I was still disappointed. It should be stable at this point.
- genghizkhan 4y agoI've never tried Ubuntu, but zfs has been rock solid on Fedora, Arch and Debian for me. Whatever the issue that hit you was, I hope recent versions of Ubuntu/zfs have fixed it.
- mkj 4y agoThat corruption was caused by a patch that Ubuntu created themselves, it was never in upstream. ZFS on any other platform would be OK. https://bugs.launchpad.net/ubuntu/+source/zfs-linux/+bug/1906476/comments/36 https://bugs.launchpad.net/ubuntu/+source/zfs-linux/+bug/190...
- kilburn 4y ago> To make use of snapshots, the backup drive gotta be Btrfs as well. The compression level was turned up to 14 this time (default was 3): Isn't this useless? My understanding is that compression is only done at file write time. When you "btrfs send" a snapshot, the data is streamed over without recompression, so there's no point in setting up a higher compression level in the backup disk.
- yakubin 4y agoIt would probably make sense to display <subdomain>.srht.site for submissions which match this domain pattern, similar to github.io sites.
- chocolatesnake 4y ago+1. This style of domain shortening should use https://publicsuffix.org/ https://publicsuffix.org/ to determine how to trim subdomains. And lo, https://publicsuffix.org/list/public_suffix_list.dat https://publicsuffix.org/list/public_suffix_list.dat contains "srht.site".
- anotherhue 4y agoPerhaps we can call this a "Considered State" system. No more haphazardly rearranging bytes on your drive. Managed OS objects, linked at boot and your home dir / config dirs under VCS. I use nixos with zfs on /home, /nix and /persist. Everything else is tmpfs, including /etc. Mostly you can configure applications to read config from /persist, but when not, a bind mount from /etc/whatever to /persist/whatever works pretty well. I will never use a computer any other way again.
- cesarb 4y ago> After freeing the new SATA SSD, I also filled it with butter. Yes, all the way, no GPT, no MBR, just Btrfs, whose subvolumes were used in place of partitions I would not recommend doing that. It might work for now, but there's a high risk of the disk being seen as "empty" (since it has no partition table) by some tool (or even parts of the motherboard firmware), which could lead to data loss. Having an MBR, either the traditional MBR or the "protective MBR" used by GPT, prevents that, since tools which do not understand that particular partition scheme or filesystem would then treat the disk as containing data of an unknown type, instead of being completely empty; and the cost is just a couple of megabytes of wasted disk space, which is a trivial amount at current disk sizes (and btrfs itself probably "wastes" more than that in space reserved for its data structures). Nowadays, I always use GPT, both because of its extra resilience (GPT has a backup copy at the end of the disk) and the MBR limits (both on partition size and the number of possible partition types).
- LanternLight83 4y agoNeat! I went looking for ways to create a protective MBR, learned a lot from the Gentoo wiki, some interesting info about how Windows does things in the link below, but the way to achieve this seem to be to just format the disk as GPT and then truncate it to the MBR (or just use one big GPT partition). https://thestarman.pcministry.com/asm/mbr/GPT.htm https://thestarman.pcministry.com/asm/mbr/GPT.htm
- cesarb 4y agoI would also not recommend formatting the disk as GPT and "truncating it to the protective MBR". Not only there's a good chance of the GPT re-appearing on its own, either because some software noticed it was corrupted and copied it from the backup copy at the end of the disk, or because some software noticed it was missing and created a new one, but also there's a chance of it once again being treated as if the whole disk was empty (since the "protective MBR" says it's a GPT disk, and the GPT has no entries). If you want to have just the single MBR sector, then create a traditional MBR with a single partition spanning the whole disk instead of a GPT "protective MBR". But that will not gain much, since you should align your partitions (IIRC, usually to multiples of 1 megabyte) for performance and reliability reasons (not as important on an HDD, where you can align to just 4096 bytes or even 512 bytes depending on the HDD model, but very important on an SSD), and the space "wasted" by that alignment is more than enough to fit the GPT.