3 ms·
Great write-up about the troubleshooting process! Regarding the exact case, there is a slightly deeper issue. XFS enqueues inode changes to the journal buffers
by gfv 4y ago
Great write-up about the troubleshooting process!
Regarding the exact case, there is a slightly deeper issue. XFS enqueues inode changes to the journal buffers twice: the mtime change is scheduled prior to the actual data being written, and the inode with the updated file size is placed in the journal buffers just after. If the drive is overloaded, the relatively tiny (just a few megs) journal buffers may overflow with mtime changes, and the file system becomes pathologically synchronous. However, since 4.1something, XFS supports the `lazytime` mounting option that delays the mtime updates until a more substantial change is written. Without it, the journal queue fills up at roughly the speed of your write() calls; with it, at the pace of the actual data hitting the disk, so even in highly congested conditions your application can write asynchronously -- that is, until dirty_ratio stops your system dead in its tracks.
- tanelpoder 4y agoAuthor of the blog post here. Thanks for the feedback and the extra XFS details (I'm not a big XFS expert). Is the XFS "lazytime" the same thing as the "relatime" mount option?
- gfv 4y agoNo, relatime applies to how often atime is updated, while lazytime controls how often all three file timestamps (atime, mtime and ctime) are written out to disk. They are orthogonal: you can have strictatime+lazytime to have accurate atime tracking that generates no disk IO on reads. The downside is, of course, if your system crashes, the non-persisted atimes will be unreliable.