4 ms·
Better title: “nother reason why my code is slow and I’m logging too much”
by dstroot 9y ago
Better title: “nother reason why my code is slow and I’m logging too much”
- deathanatos 9y agoYeah. The whole article seems to be a bug in the logging library that's in use here, and an area of contention in the kernel? (How many times is his application logging, to cause that much contention?) I'm not sure how Docker figures into it, aside from it easily lets one run multiple instances of an app on the same hardware, but stuff like supervisord will also do that? I'm not entirely clear as to why a logging library needs to call fadvise; a log file, is, I presume opened in append-only mode. Isn't "append" sufficient advice to the kernel? Also, fadvise needs byte ranges, and I have no idea what you'd pass for a log file…
- asdbffg 9y agoUsing O_APPEND does not imply, that kernel needs to purge the pages from cache ASAP, does it? Removing pages from cache may be expensive operation by itself, so I presume, that it is avoided by default. More importantly, if the disk can not catch up, the log data is going to end up waiting in page cache anyway (typical case of bufferbloat). Linux kernel does not have telepathic abilities to balance needs of crazy logger and other applications in system, so without resolving underlying issue (bufferbloat), those writes would take up too much cache, potentially bringing down disk performance of other applications. fadvise() may schedule quicker eviction, effectively acting as syscall version of vm.dirty_ratio. Of cause, that does not resolve the problem, — just moves it to different layer. The real solution is either 1) blocking the apps until their logs are fully written (for example, by using O_DIRECT) 2) showing those apps middle finger and throwing away some of their logs (AFAIK, this is occasionally done by syslog).
- pacavaca 9y agoThe point was not to complain about how bad the Docker is but rather to highlight that a lot of unexpected things may come from the fact that the kernel is shared. "Logging too much" was just a reason for the posix_fadvise being called too often but this becomes a problem ONLY when the kernel is shared. In case of virtualization, everyone gets its own version of fadvise (the kernel) and the conflict doesn't happen.
- cjhanks 9y agoIf your VM call to `fadvise` is not calling the underlying host kernel operation, is it even working?
- pacavaca 9y agoHmm. I'm definitely not an expert in hypervisor implementations but I would guess that it should not proxy any calls to the host kernel...
- dullgiulio 9y agoBut as fadvise can reliably only be implemented in the host kernel, the one in hypervisor is basically a noop.
- fulafel 9y agoIn the case under discussion, the dontneed fadvise would tell the guest kernel page cache to discard written data after write. So it is useful without the host kernel knowing about it.
- asdbffg 9y agoVM certainly does "call underlying host kernel operation", it just does so indirectly — the guest userspace calls fadvise(), kernel implementation of fadvise() asks the virtio disk driver to perform particular read/writes, the virtio driver asks underlying kernel disk driver to read/write individual disk sectors (without knowing, that they are related to specific file in guest filesystem). This specific bug was caused by putting high load on "kernel dentry cache", e.g. a contention for memory structure, present in kernel memory. Guests normally don't share memory, so contending for it was avoided. Incidentally, there are situations, when different guests can compete for same memory — when VM uses so-called "memory deduplication" techniques. Which is why enabling that stuff on production systems may be a bad idea.
- cjhanks 9y ago