19 ms·
Single log line is 49KB+ (ext4) / 110KB+ (btrfs) of systemd-journald disk writes
- Storewide 2mo ago[flagged]
- ValdikSS 2mo agohttps://github.com/systemd/systemd/issues/40262#issuecomment-5169482357 https://github.com/systemd/systemd/issues/40262#issuecomment...
- davidricodias 2mo agoThanks for that comment. Out of curiosity how did you come up with that setup? For me most of that test suite sounds alien
- ValdikSS 2mo agoIf it weren't mmaped files, I would use strace/gdb, or even fuse proxy file system. But these are mmaped, I don't know any easy debugging or monitoring solution besides writing kprobes/systemtap hooks. How would you debug it?
- smartmic 2mo agoI recently looked into disk usage of journald and was also shocked. My next step towards peace of mind is https://www.devuan.org/os/init-freedom https://www.devuan.org/os/init-freedom Will try it out as next distro for my Debian system, longtime experience with Void Linux (runit) on another box is great.
- ValdikSS 2mo agoMany applications hammer the disk even if the developers don't believe this is an issue, not only journald, unfortunately. It's my third attempt to make my regular Linux desktop less disk-chatty. This is a huge issue for btrfs and for COW FS in general, because they have massive write amplification for small and frequent writes (38,7 TB written to my idle desktop SSD in 2 years). If you're interested, here are my findings this time so far: - workrave: 60 second stat sync https://github.com/rcaelers/workrave/pull/717 - kde klipper: saves to disk on every copy, even if permanent storage is disabled https://bugs.kde.org/show_bug.cgi?id=501030 - kde plasmashell: saves qt shader cache each time notification popup disappears https://bugs.kde.org/show_bug.cgi?id=523805 - bitwarden firefox extension: tries to connect to desktop application every 10 seconds, writes about every failure to browser's WebStorage 14+ KB https://github.com/bitwarden/clients/issues/22192 - firefox datareporting/glean: very chatty .mozilla/firefox/xxx/datareporting/glean/db/data.safe - ipfs: writes every received DHT announce to disk, 20 GB in 3 hours https://discuss.ipfs.tech/t/constant-writes-to-datastore-log/20316 - mailcow: redis saves data every 5 minutes https://github.com/mailcow/mailcow-dockerized/pull/7405
- doublepg23 2mo agoThe two I most often see in Ubuntu's dmesg are: audit - appears to be some sort of AppArmor logging? br[] - bridge interface docker uses consistently rebuilds itself? May be related to docker compose networking.
- marginalia_nu 2mo agoYeah docker does a ton of network stuff when you start/stop containers, depending on your configuration. It's extra fun because it can drop existing connections when that happens. Had a process quietly in a crash loop for a solid month on my workstation until I figured out what was causing my random network outages.
- graemep 2mo agoI noticed plasmashell is write heavy and logs to journald a lot so I have just switched to XFCE partly for that reason.
- 3abiton 2mo agoUnfortunately it's not always easy to move away from systemd. I still run void on one of my machines, but aur make things so much easier.
- otterley 2mo ago> This is a huge issue for btrfs and for COW FS in general, because they have massive write amplification for small and frequent writes Have you considered using a different fstype like XFS for this? btrfs is good for homedirs, but I wouldn't necessarily use it for other filesystems (/usr, /var, etc.)
- ValdikSS 2mo agoI don't have experience with xfs or zfs, and I don't have a free stand for experiments right now unfortunately.
- michaelmrose 2mo ago
- p_l 2mo agosystemd-journald has one of the most deranged log file formats I have ever dealt with, and one of the worse user interfaces, too. I am not again binary logs, or logs in a database. It's just yet another time I deal with good ideas implemented horribly, horribly badly when it comes to systemd.
- hedora 2mo agoI’ve been using devuan more or less since day one. I highly recommend it. FreeBSD isn’t too shabby these days either.
- rustcleaner 2mo agoI wish Qubes Domain-0 was a customized Gentoo with OpenRC. Fedora with systemd was a poor choice to base off. Nobody should have let Poettering have the influence he was given over userland, systemd is an almost irrevocable mistake.
- tryauuum 2mo agohello ValdikSS! nice to see you alive
- amluto 2mo agoOoh, mmapped writes. I make that mistake once, years ago. :) I posted a comment in that GH issue.
- zbentley 2mo agoSay more? Sounds like a good story
- speed_spread 2mo agoMy caveman understanding is that mmap writes are bad for transactional accesses because you have little to no control over sync. The OS can decide to commit changes to disk anytime, in any order which is the opposite of what you want for anything ressembling a database.
- zbentley 2mo agoBy default yes, that’s true. But while there isn’t a reliable don’t-flush-this-page system, there definitely are ways to force the flush of specific ranges in an mmapped file. But you’re generally right. I think that’s why most databases have the notion of a WAL, which is carefully append-only. But the non-WAL data files in most DBs I’ve used are accessed via mmap.
- amluto 2mo agoThe problems are much more than just sync. A long time ago I had this over-optimistic idea: x86 hardware (and probably most other hardware) has these cool hardware-managed dirty page bits. So you would write to a mapped page, not even take a page fault, and the hardware would record that it's dirty. Later on the kernel would notice and flush. Excellent performance. Hahaha. It's much much much more complex. For various reasons (maybe good, maybe bad -- see below), Linux barely uses the real hardware dirty bit. Instead, when you map a file as shared-writable, at first it might not really be mapped at all. If you read it, it gets faulted in and becomes readable. When you first try to write to it, a page fault is generated, and, on non-FRED x86, the page fault itself is very slow. The kernel will do things, including calling into the FS and updating atime [0], to make the page logically writable. It updates the page tables so that the CPU knows it's writable, and it sets the dirty bit right then (after all, this is a bit faster than letting the CPU set it immediately thereafter when you retry the faulting write). Okay, now it's writable. Writes are essentially free until the kernel decides to write the data back to the disk. The kernel will mark the page non-writable (because it wants to get notified the next time you try to write to it) and flush the TLB (which is extremely expensive, especially on x86 systems that aren't the latest AMD CPUs). And it will write the page back, more or less as if you had used normal syscalls to write it. There's more fun, though. Some filesystems and/or backing stores need "stable pages" -- they need the page cache pages that are being written to not be modified while being written back. btrfs, for example, wants to checksum the data and then write the data and the checksum out consistently, and if something changes the data while it's being DMAed, then this can't happen. So special locks might be taken to delay future writes to the page until writeback is done, and that includes blocking the "make writable" page fault handler. Oops, there goes performance. Could the kernel do better? Probably. Will it? Unlikely in the near future. I've contemplated a special mechanism to map a "fast write" window onto a file that would be permanently writable and use the hardware dirty bit to tell the kernel when to transfer the data out. Even if anyone ever implemented this, it would be a very specialized thing, it would incur polling overhead, and it would be utterly silly to use it for something like syslog. Just use pwrite or io_uring unless you have actual evidence that mmap is better. mmap read is a different story, of course. [0] I think that updating atime at make-writable time instead of at writeback time is both non-performant and semantically incorrect. I've never convinced the maintainers well enough, though.
- barrkel 2mo agojournald is IMO the worst part of the systemd ecosystem. You're better off using it only as a router and not storing any logs in it. The indexing system it uses is slow and provides no control over chatty subsystems - you cannot truncate the logs for just a single identifier. For all the use indexing is doing you will get better performance out of a modern grep like ag or rg. Structure is worth something but it's better off somewhere other than journald.
- graemep 2mo agoI recently put a lot of effort into reducing logging because of excessive writes. It was so much easier when everything had its own log and you could just look at which files were growing.
- e2le 2mo agoI would much rather that they had used an existing database file format. Sqlite3 is robust and already present in the default installation of most Linux distributions. Querying system logs with SQL would be cool and likely faster than using the sd_journal API with all it's weird quirks.
- deleted 2mo ago[deleted]
- Walf 2mo agoText or text-like (e.g. text content with simple control char delimiters for metadata) would be far superior than the slow-down from Sqlite's safety mechanisms. Optimising logs for read, at the expense of write, is a bad pattern to me.
- ahartmetz 2mo agoRead optimized? That is funny because reading logs from journald is dog slow compared to, you know, log files.
- 2mo ago
- sam_lowry_ 2mo agoCool to see @ValdikSS here as well. The guy never sleeps or he is AI in disguise ;-)
- mono442 2mo agojournald has never been of great quality. It somehow manages to be visibly slower than grepping gzipped text logs.
- pengaru 2mo agoit has bad scaling properties esp. if you have many journal files the last time I contributed to journald upstream was to fix a degenerate behavior with many journal files: https://github.com/systemd/systemd/commit/176f73272e6e3116caab3900eb553be54f520a68 https://github.com/systemd/systemd/commit/176f73272e6e3116ca... that makes a dramatic difference for those hitting this case, but it only gets things from nearly unusable to slow-as-usual.
- pudgywalsh 2mo agoHow do you try to copy Windows NT's Event Log — which is essentially unchanged from the 1990s when systems ran on 32MB of RAM or less — and fail so spectacularly? The first thing I do on a Linux system is install a proper syslog daemon.
- rasz 2mo agoOne of the first things I do on win10 is disable most of excess logging.
- breakingcups 2mo agoBut why, though? I have never seen a performance hit from it that would warrant that.
- rasz 2mo agoDont like SSD wear for no reason. I wont look at those logs anyway on my personal gaming pc so its all useless.
- pudgywalsh 2mo agoYour issue is greatly exaggerated. Enterprise users increase the logging and I've never heard of premature SSD failure due to this. The event log is capped in size (adjustable). It's nominally < 100MB. Your games continually dumping GBs of data into local cache on the other hand...
- rasz 2mo agoSize cap doesnt mean much when its a constant stream of small writes.
- sidewndr46 2mo agoHow do you do that?
- pengaru 2mo agoI'm probably the main person responsible for making journald usable at all. But I never really made any effort to change the on-disk structure or how writes were performed. My focus was more on the read performance for journalctl and stability of the daemon. Back when I was paid to fix things in journald at CoreOS ages ago, it couldn't even avoid getting killed by its own service watchdog. My impression back then was the on-disk format dispersed the information too much within the same file, and those individual datums being written at discontiguous offsets were quite small, far smaller than an IO block size or even a disk sector size. Seemed like a write amplification problem due to the file format. If you write a few bytes into some arbitrary position within a file, the storage has to write back the whole block, despite your only changing a tiny fraction of it. If those few bytes happened to cross a block boundary, guess what? two blocks get written. The format had no consideration for these block-oriented storage details, then doing the IO via mmap rubs salt into the wound since the kernel has to try guess what to prefetch asynchronously... but I don't think that aspect amplifies the writes above what plain buffered IO would do - maybe I'm wrong. I'd expect the mmap aspect to be causing more/mispredicted reads, and polluting the page cache with unrelated contents (you tend to end up with the entire journal cached IIRC, if you have enough memory). I suppose there's probably compounding of the write amplification problem since the kernel will be dirtying pages at page size granularity vs. 512b sectors, and you have the same issue of small writes landing on page boundaries dirtying two pages. So that aspect of using mmap for the writes probably is exacerbating the problem.
- ValdikSS 2mo agojournald uses hash tables, I think it update it on every new log line, although I didn't debug it in depth yet. https://github.com/systemd/systemd/blob/199f75205b9c0625bf56e229b12b95a645ed7a6c/src/journal/journal-file.c#L1083 https://github.com/systemd/systemd/blob/199f75205b9c0625bf56...
- pengaru 2mo agoThe format is documented https://github.com/systemd/systemd/blob/main/docs/JOURNAL_FILE_FORMAT.md https://github.com/systemd/systemd/blob/main/docs/JOURNAL_FI...
- skullone 2mo ago[flagged]
- rasz 2mo agoOh how I love totally predictable poetterings reply to previous bug report that got closed because "measuring it wrong" and "this is not a support forum".
- otterley 2mo agoThis issue report feels like it ought to be accompanied by a fix. If you think you can do better than journald's existing format, propose a new one with tests to prove it. GenAI makes this much easier than it used to be.
- lucb1e 2mo agoOr you talk with the others first to see what kind of setup everyone thinks is good. I'd find it strange if someone barges into my project with a pull request that fundamentally changes the design of a major component
- otterley 2mo agoSure, a concrete proposal first would be a good idea. That said, would you look a gift horse in the mouth?
- ericpruitt 2mo agoBecause you become responsible for feeding and taking said care of horse and dealing with any technical debt associated with it. If someone submits code to a project that I maintain that's going to make my life difficult in the future, I'm not going to accept it.
- Brian_K_White 2mo agoIt doesn't matter how free a turd sandwich is.
- otterley 2mo agoThere's no proposal that we can evaluate to determine whether it's a turd sandwich or not.
- shawnz 2mo agoDesigning a new on-disk format seems like a pretty far reaching architectural decision... I don't think that's an appropriate target for a drive-by fix from a new contributor
- quotemstr 2mo agoSystemd should just use DuckDB. It's perfect for this job. "But isn't it an OLAP database? Shouldn't you use SQLite for something that's vaguely real-time?" Eh, in this instance, I think I'd prefer the columnar design and automatic compression DuckDB affords. Log entries have lots of little fields, many of which are unchanging from row-to-row, and DuckDB excels at storing this kind of data. BTW: no, you don't need O(N*log(N) writes for DuckDB. No, you're not doing a whole block-group write for every message. No, Parquet is not a magical solution. I mean, maybe it's fine, but DuckDB is already columnar, and arguably better at it. Seems like there are a lot of mistaken impressions about DB storage engines out there.
- marginalia_nu 2mo agoParquet is probably an even better option. Columnar, compression, fast, succinct. All good things. You can read them with DuckDB, but you don't end up with O(log n) writes -- which is, to speak plain English, batshit fucking insane for a system logger. What those cursed writes buys you is O(log n) reads, but there's just no scenario that is necessary. If you have literally any time or subsystem constraints, parquet's predicate pushdowns means you get plenty fast access even with a full scan.
- orf 2mo agoNo, not at all. Parquet is great for building static content incrementally, but it’s not great for this: the aim is durable writes (it’s a log system after all), but with parquet you need large row group batches. Worst case (low log volumes and a time-based flush) you’d end up with loads of tiny row groups. You also need metadata in the file footer, so you can’t query it until the file is “done”. When is that?
- lokar 2mo agoFor 99% of installs the basic assumption that local logging (with local reading) is the primary mode is just wrong.
- zbentley 2mo agoWhat do you mean? That has described the vast majority of Linux systems I have ever touched, professionally or personally. Even corporate environments with log aggregation tail system logs rather than having them directly shipped elsewhere. The rare exceptions to this are some embedded devices without much durable storage, or tightly regulated environments in which log data is considered radioactive.
- hedora 2mo agoI’ve met people that think the log should be remotely stored and not written locally, since it’ll be shipped to splunk or whatever anyway. Those people change their minds the first time a machine has intermittent network issues, and the logs needed to debug it are lost (or worse, the log buffer fills, then stdout fills, which backpressures the application, creating an outage while simultaneously eating the logs).
- deleted 2mo ago[deleted]
- lokar 2mo agoBy count, most installs will be the large cloud providers
- zbentley 2mo agoAgreed. And in my experience , most large cloud providers’ Linux systems I’ve worked on (either their VMs as a tenant or their underlying hardware as an employee) log locally and ship additionally.
- 2mo ago
- d3Xt3r 2mo agoOkay, so how do I disable journald and switch to something else, without getting rid of systemd completely?
- marginalia_nu 2mo agohttps://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html https://www.freedesktop.org/software/systemd/man/latest/syst... See StandardOutput= and StandardError=.
- d3Xt3r 2mo ago> Oh no! Bad Request > Error: access denied: error in challenge meta-refresh: mismatched token God I hate the modern web. I get that anti-bot measures are necessary, but at what cost?
- marginalia_nu 2mo agoYou get the exact same information in $ man systemd.exec Though reading the question again, I should have probably linked to the equivalent of $ man systemd-system.conf as well, that's where you can set the default behavior across systemd, not per-service as the first man page is.
- pineapplepizza6 2mo ago[dead]
- deleted 2mo ago[deleted]
- CrimsonCape 2mo agoOk, so I should set these to null? I saw elsewhere someone set journald storage to volatile. Which of these approaches is better?
- vachina 2mo ago
- otterley 2mo agoSomething must have happened along the way, because this was not the original design intent of the database (emphasis mine): """ The native journal file format is inspired by classic log files as well as git repositories. It is designed in a way that log data is only attached at the end (in order to ensure robustness and atomicity with mmap()-based access), with some meta data changes in the header to reference the new additions. The fields, an entry consists off, are stored as individual objects in the journal file, which are then referenced by all entries, which need them. This saves substantial disk space since journal entries are usually highly repetitive (think: every local message will include the same _HOSTNAME= and _MACHINE_ID= field). Data fields are compressed in order to save disk space. The net effect is that even though substantially more meta data is logged by the journal than by classic syslog the disk footprint does not immediately reflect that. """ See https://docs.google.com/document/u/0/d/1IC9yOXj7j6cdLLxWEBAGRL6wl97tFxgjLUEHIX3MSTs/pub https://docs.google.com/document/u/0/d/1IC9yOXj7j6cdLLxWEBAG...
- simoncion 2mo agoIf it ever worked like that, then gradual accretion of (mis)features and misguided enhancements pretty clearly broke it. Based on my years and years and years of reading about and using the output of the Systemd Project, there's really clearly no Linus Torvalds on the project to hold the line on software quality. Edit: Looks like someone who did a ton of work attempting to get journald even vaguely usable has chipped in with additional information. [0] My hunch is that the current set of people working on the Systemd Project are going to be supremely disinterested in fixing the problem... and might even be entirely unable to fix it. A project this large and sprawling that runs for this long without a solid commitment to quality doesn't tend to retain many very highly-skilled individuals. [0] <https://news.ycombinator.com/item?id=49291376 https://news.ycombinator.com/item?id=49291376>
- throw0101a 2mo agoI was always curious why they created their own format rather than leveraging SQLite, OpenLDAP's LMDB, etc.
- jck86 2mo agoThe cherry on the cake is that you practically cannot filter journald. The only option is limiting by severity (e.g. errors and higher) or switch to non persistent journald storage and forward to rsyslog and filter there. Am a bit vague on the details but sometimes a driver goes bezerk and starts logging many times per second, e.g. a bug in amdgpu after resume from suspend. Took a while to get that filtered which luckily was only possible because it were kernel messages (dmesg), but for a while I had to disae persistent kernel logging which is dat from ideal. I get that for certain core parts simplicity is more important than features. But journald is just too basic to enable persistent storage but I also don't want to switch it off.
- mzajc 2mo agoIf you have systemd>=253 you can make use of LogFilterPatterns[0] (in .service files), but it's really unpredictable, cumbersome to work with, and does not work with user services or non-service log sources. [0]: https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#LogFilterPatterns= https://www.freedesktop.org/software/systemd/man/latest/syst...
- jck86 2mo agoThanks for this, was not aware this was added. Though better would be to have a global option for this, and a tunable whether to filter this completely or only for persistent storage. For the last few days I have been monitoring journald with iotop and found in my case storage use was not excessive at the moment. And there are rate limit options, but I'd really like there to be system-wide filtering options. For servers I rarely see journald persistence enabled while it is actually very valuable for debugging crashes and other issues. Way easier than regular log files. Though also more fragile and more difficult, so improvements are very welcome.
- sidewndr46 2mo agoyears ago, I set Storage=volatile on almost all the journalD configurations I have. This largely solved this kind of problem.
- itvision 2mo agoGood for your personal devices, very not good for servers where you need to have any sort of accountability and security trail.
- otterley 2mo agoThat’s why you ship the logs off host to a central collector.
- itvision 2mo agoNice in theory in practice you always retain them locally as well, just in case the network connection goes down. If network egress fails and logs are pushed in real time over a connection with no local backing, you face an ugly tradeoff: either drop log data silently (loss of visibility during the very network partition you need to debug) or apply backpressure to services (potentially hanging applications when logging buffers saturate).
- otterley 2mo agoAgreed. That's what the log agent's disk buffer is for. That can still be used even if the journal itself is on volatile storage.
- hedora 2mo agoSomeone should implement a new operating system that can efficiently handle text processing. It could have some simple tools that let you generate reports, display them on screen, and compose tools for that sort of thing in a natural way. We could call it UNIX.
- 0x_rs 2mo agojournald is awful for many reasons, but what makes it worse is that everything running on your machine thinks it has any rights to dump all the logs it wants unprompted. Open a file picker and kio will decide it's a good idea to spam tens or hundreds of thousands of entries into it a day, listing every single file you have in a directory with some log such as "No node found for item that was just removed" and that has zero impact to the user whatsoever. You almost need to keep a script tracking all the journal floods for every new service to make sure it's not treating your system log as its dumping ground. To be fair, the kernel and usb peripherals can also have a bad day and spam 3 million lines an hour into it, think input irq status -75. It's too much of a chore to keep up with all the program-level configs (if they have them) and service files, but LogFilterPatterns in systemd can help in an unintended way: you can make one log blacklist with a .conf file in /etc/systemd/system/service.d/, and put in there all the patterns that spam your journal one by one, don't even have to chase misattributed loglevels. It just looks something like: [Service] LogFilterPatterns=~I am a completely useless log entry LogFilterPatterns=~I am another useless log entry But it doesn't pick up on identifiers and doesn't do anything for kernel spam. It's only great to make some messages shut up. Also, I'd consider any btrfs install that does not have nocow on cache, journal etc. to be defective.
- greatgib 2mo agoThat was the task for years for syslog services that dealt with it without issue.
- touisteur 2mo agorsyslog is an incredible piece of software. Every time I'm looking for something to do with logs, opening the docs or googling finds the feature for me and myriads alternatives. I know it still exists and use it heavily on any system I'm in charge of, but there's some regret at having a dual system with journalctl...
- giov4 2mo agoI can confirm this, used it for many years, 3 keywords: efficient, reliable, useful. all 3 missing on journald,in my experience i saw it inefficient also on configuration level, unreliable because of loosing loglines on crash or reboot and not useful since to look at logs i need 3 commands, verbose parameters and 5 google search to find them. syslog experience? very efficient also on heavy load production instances, never lost a log, pipe grep and jq and you have the info you need. so what I experienced is that a default linux install was shipping a rock solid logging system by default, reliable and usable and everybody knew what was where and you will find it. now i just have fancy stuff, units etc and lost all of that. no I dont need to tune config parameters on a default install to have working basic logging tnx.
- greatgib 2mo agoSystemd things being horse shit as usual because it was vibecoded even before LLM existed. And there are still people that said that systemd and tools are awesome because they never encountered any of the countless ridicule bugs.
- micw 2mo agoWhat does the (currently latest, https://github.com/systemd/systemd/issues/40262#issuecomment-5288996192 https://github.com/systemd/systemd/issues/40262#issuecomment...) comment mean? Who are the "large folio people"?
- pseudalopex 2mo agohttps://news.ycombinator.com/item?id=44113020 https://news.ycombinator.com/item?id=44113020
- micw 2mo agoSo is this a regression or a bug in systemd triggered by this change?
- 1blackeagle1 2mo ago[dead]
- adrian_b 2mo agoI completely agree with one of the comments from there: > But the fundamental conclusion is: the design was wrong. It should not have used mmapped writes. pwrite would have been far better. It really does not make any sense to use memory-mapped files when writing logs. Not even pwrite makes sense, because logs should normally be written by opening and using the log files as append-only sequential files. Only when reading logs, to search for problems, accessing them as read-only memory-mapped files is OK. Actually not only for logs, but almost always, read-write memory-mapped files are either inefficient or too complex to use (i.e. to avoid problems you must carefully use msync and/or madvise, which eliminates the simplicity that makes memory-mapped files preferable to using pread/pwrite). It is better to use memory-mapped files only for read-only accesses, using the appropriate option flags in open and mmap.
- giov4 2mo agogreat and clear summary thank you! I would laso add that if my design decisions or development actions lead to an issue affecting multiple linux distro defaults I would feel responsible and rush for a solid fix instead of this https://github.com/systemd/systemd/issues/15292#issuecomment-777276876 https://github.com/systemd/systemd/issues/15292#issuecomment...
- giov4 2mo agoadditionally and ironically in a case like this AI would have been probably more efficient and already resolved the problem with a PR cycle instead of human histeric gate keeping.
- smartmic 2mo agoThis is really astonishing. Are there no checks and balances in place for design decision in such a critical system component? What were the thoughts of all the major distros when they decided to go with systemd then?
- pineapplepizza6 2mo ago[dead]
- microgpt2 2mo ago[dead]
- acrush 2mo ago[flagged]
- zbentley 2mo agoMy hunch having looked at the journald code as an amateur is that this write amplification is coming from scattering, with a few possible sources: 1. Writes try to compress away duplicate metadata at the application layer, which causes them to issue scattered writes when new metadata shows up. 2. Indexing is also surprisingly log-line/application-layer aware, such that index writes might also be scattering. 3. The indexes themselves seem like they could benefit from an append-mostly write model with periodic compaction rather than a mutate-in-place model. 4. I was surprised that the journal’s “WAL” doesn’t seem to be a major concern of a lot of the code. For a database, supporting reads “through” the WAL with periodic application back to the data files (“checkpoints” in RDBMS) seems like something I’d expect to see more of here. But I don’t really have deep understanding of the code, so I may be missing that it’s doing that already. The choice of mmap instead of regular file writes here isn’t, as others have proposed, a design flaw. I think that makes sense given what journald is (a database) and how significant its durability concerns are. And it looks like the code does spend a lot of time trying to be careful about which blocks/pages are dirtied. But this is a famously hard-to-get-write (ha!) area so perhaps defects are present at that layer. The systemd developers are talented in their area; I am not a systemd hater. However, “talented at low-level OS design” is not the same as “talented at building a database from scratch”, and I think that shows here. I strongly feel like this system could be a wrapper around SQLite, which is definitely something that could be integrated everywhere journald is used (license-wise and compatibility-wise). I’m puzzled as to why that wasn’t chosen as an approach: a SQLite vfs implementation that handled compression and online rotation seems like it would have resulted in a design that’s both more interoperable and less prone to flaws like this one. I also think that a per-log-emitter setting that doesn’t eagerly persist to disk (wait for page cache flush) would be very useful to have available—perhaps even as a default—for user-level/init6 level logs that are OK with a potential for data loss on kernel panic.
- itvision 2mo agoFeatured on: https://northeasttimes.com/2026/08/14/a-single-log-line-eats-49kb-of-disk-and-linux-admins-are-furious/ https://northeasttimes.com/2026/08/14/a-single-log-line-eats...