3 ms·
I have been running k8s clusters at utilizations far beyond 50% (up to 90% during incidents). For web services/microservices, so tail latencies were important.
by treffer 5y ago
I have been running k8s clusters at utilizations far beyond 50% (up to 90% during incidents). For web services/microservices, so tail latencies were important.
The way we solved this? 1. Kernel settings. Check e.g. the settings of the Ubuntu low latency kernel for example. 2. CFS tuning. Short timeslices. There are good documentations on how to do that 3. CPU pressure. We cordoned and load shedded overloaded nodes (k8s-pressurecooker).
By limiting the maximum CPU pressure to 20% you can say "every service will get all the CPU it needs at least 80% of the time on most nodes". This is what you want. A low chance of seeing CPU exhaustion. This is needed for predictable and stable tail latencies.
There are a few more knobs. E.g. scale services such that you use at least one core as requests are effectively limits under congestions and you can't get half a core continuously.
Very nice to see that people go public about this. We need to drop the footprint of services. It is straight up wasted money and CO2.