5 ms·
One source of high load average spikes that I've seen in my job is when a process crashes and generates a core dump. While the core dump is being written, all t
by sreque 9y ago
One source of high load average spikes that I've seen in my job is when a process crashes and generates a core dump. While the core dump is being written, all threads in the process are in the TASK_UNINTERRUPTIBLE state even though they are doing absolutely nothing, and as such they all count towards the load average as if they were spinning on on a CPU core. If the total virtual memory of the process is large, say in the multi-GB range, coredumping can take on the order of a minute, and Linux will report an unreasonably high load average if that process had a lot of running threads.
Things like the above scenario make me treat the load average metric with a lot of skepticism. I would much rather use other metrics to infer load.
- lotyrin 9y agoI rarely recommend alerting monitoring or any kind of action based on load averages or more generally any metric derived from queue lengths. It's trends in high-quantile queue latencies your users (and therefore you should) care about.
- haimez 9y agoKind of ironic, given that the whole article is about the divergence of system load from being a queue length metric.