4 ms·
What’s your reasoning? From my experience, a memory limit, if hit, will cause periodic OOM and restart; and a CPU limit will cause slowing-down of the processe
by bboreham 8y ago
What’s your reasoning?
From my experience, a memory limit, if hit, will cause periodic OOM and restart; and a CPU limit will cause slowing-down of the processes.
The tale already mentions restarting and slowing-down; it’s not clear to me that adding more would have improved anything.
- jpgvm 8y agoThe issue was the Docker daemon was affected due to poor resource configuration. It should run at a higher priority and with guaranteed resources so it can't be impacted by user mode tasks. k8s has been slow to catch up in this area but finally has priorities and preemption. That said Docker normally doesn't run in a pod so it's also a matter of setting the node total allocatable CPU to something reasonable and ensuring all pods spawned by kubelet are nested under a namespace that has a lower priority than the system tasks (kubelet itself, Docker if you use it etc).
- chupasaurus 8y agoAnd it could be easily done by adjusting Docker's systemd service unit.
- merb 8y agoyeah after reading that I came to the same conclusion and tought that he mostly blames it on cluster size, which is just wrong.
- bboreham 8y agoKubernetes has a “-system-reserved” flag for that purpose. Limits are something else entirely.
- KaiserPro 8y agoYou want to signal early that a node is unable to do what is requested of it. Its the same with unbounded queues, you need to put in limits so that alarms are triggered much earlier. infinitely spinning up new things is generally bad. Unless you are responding to an external signal. Minimum deploy times again are useful. There is a financial cost (on the cloud, not on real tin) to short term instances. So you need to have them on as long as possible.
- bboreham 8y agoSounds like a Kubernetes resource request, rather than a limit.
- tedk-42 8y agoif you've got an app that constantly uses more memory then it's either a memory leak or poor tuning which results in that. An app without a CPU limit can completely consume the hosts CPU and take down everything else on that node. If you've got a very large instance with a lot of apps, a bunch of them are gonna experience a poor QoS. Their tale suggests their sidecar pods/containers (which likely didn't have limits) made their production issues worse. a CPU limit on those pods would have limited the damage those pods caused to the rest of their infra.