5 ms·
> We also considered simply abandoning autoscaling altogether and pinning to a calculated value, but this would hide performance regressions in the code by abso
by mrbrowning 9y ago
> We also considered simply abandoning autoscaling altogether and pinning to a calculated value, but this would hide performance regressions in the code by absorbing them into a potentially enormous buffer intended for regional evacuation absorption.
I'm probably missing something here, but why wouldn't performance regressions still be detectable via utilization metrics? I understand the difficulty of determining resource allocation a priori, but I'm not sure how this relates.
- aaronblohowiak 9y agoWhat you propose would catch many changes by looking at things like cpu and latency, but the space of potential hidden resource constraints is vast enough that without empirical verifician, we cannot have high confidence that performance regressions have not occurred (eg: lock contention can be not a problem at all until it is a huge problem...)
- mrbrowning 9y agoThat’s a good point. Out of curiosity, how do those constraints get surfaced with hosts running in ASGs?
- aaronblohowiak 9y agoas the new version gets deployed and starts to take more traffic it doesnt keep up and is scaled up accordingly which establishes a new performance curve for the failover prediction system.