3 ms·
One challenge with running PostgreSQL in production on Kubernetes is that it does not behave nicely under high memory pressure. PostgreSQL likes to get NULL bac
by mslot 3y ago
One challenge with running PostgreSQL in production on Kubernetes is that it does not behave nicely under high memory pressure. PostgreSQL likes to get NULL back from malloc such that it can gracefully abort transactions (overcommit_memory = 2 in Linux). However, this model of tracking and dealing with memory pressure is still not available at the cgroup level. Instead, processes will be OOM killed. In PostgreSQL, that triggers potentially lengthy crash recovery and therefore downtime. On the flip side, the cost of deploying a hot standby is often lower on Kubernetes, but failing over due to high memory pressure can also lead to pathological behaviour.
A good discussion of this issue by Joe Conway:
https://www.crunchydata.com/blog/deep-postgresql-thoughts-the-linux-assassin https://www.crunchydata.com/blog/deep-postgresql-thoughts-th...
Hence, not completely the same as running on a VM.
- alexeldeib 3y agoThe article mentions avoiding overcommit and oom score adjust. You can avoid overcommit by always specifying requests == limits and can use priority class for oom score adjust. There are definitely improvements like memory qos/pressure handling, but not sure what you mean about those specifics, they can be handled. You can always oom the whole node, I don’t think(?) the fact that there’s a non root oom matters? So you could arguably set no pod limit, set priority class critical, and let another pod in the workload cgroup get killed.