3 ms·
They were using high-resource nodes and made the mistake of shifting to low-resource nodes for fine-grained control over the total amount of resources provision
by edaemon 7y ago
They were using high-resource nodes and made the mistake of shifting to low-resource nodes for fine-grained control over the total amount of resources provisioned -- i.e. a single 96-unit node vs 96 1-unit nodes. That was a problem because the low-resource nodes were much less efficient at processing the actual load; a portion of the resources for each node were allocated to k8s system processes, but also any idle pod consumed a higher proportion of the node's resources to do nothing. As a result the autoscaling functions provisioned even more resources than they were using with the high-resource nodes. The solution was to use medium-resource nodes that offered somewhat granular control but made efficient use of the available resources.