3 ms·
In an era before cloud, one of the major project risks was under-provisioning hardware. Load testing made sense then. It makes less sense now. I've given up on
by tkahnoski 6y ago
In an era before cloud, one of the major project risks was under-provisioning hardware. Load testing made sense then. It makes less sense now.
I've given up on load testing all together. Most applications or services quickly grow to complex to maintain in a cost-effective way with the pace of change demanded of them.
Instead every team I talk to about load testing or performance I shift it to observability. If the team can't understand current performance and load in production, any sort of load testing results in another environment will be poorly understood and hold little value.
This approach positions the team much better to react to regressions in production vs holding up work trying to create or pass a load test.
The exception I make for this is load testing for validating technology choices as an effort in risk mitigation that the technology can't perform. (i.e. Will this query work moving from SQL to ElasticSearch? What happens if I have 100x amount of data in that table?) Targetted specific scenarios, to confirm behavior of things too expensive to do in production.
I'm sure there are a few performance critical apps that need these tools, but the vast majority of software doesn't. Don't burn 100s of hours like I did to validate performance before release. Start with gaining a deep understanding of your production behavior and push for production experimentation. It is significantly less time-consuming and pays way more dividends.
- vinay_ys 6y agoPrioritising observability and load stress mitigation mechanisms/playbooks is the right thing to do for modern cloud native application service clusters. In such an environment, you could fairly easily make it possible to subject your prod deployment to synthetically generated end-to-end scenario traffic in a safe manner (through request tainting) so that you can create controlled stress situations and see how your application behaves at breaking conditions. This is important for understanding if the observability and mitigation mechanisms/playbooks are in good shape or not. It is especially important to be able to do this periodically because the functionality is continuously evolving. Cloud can lull you into a false sense of safety – not everything in cloud scales or fails over the way you expect them to at all operating points. So it is important to continuously do controlled assessments at loads 2x higher than your current peak 6 months before that becomes your reality. Of course there is always live and learn by fire approach where you wait until the traffic growth (from that frenzy flash sale your business team sprung on you) topples your stack over.
- tkahnoski 6y agoThat is a fair point. Most of my experience is in modest growth and not exponential growth companies where scaling points are less predictable.