3 ms·
Thanks for the awesome writeup. Since you opened up questions: From the start of your cutover (when you started seeing the 500s) to resolution of the various i
by firebones 10y ago
Thanks for the awesome writeup. Since you opened up questions:
From the start of your cutover (when you started seeing the 500s) to resolution of the various issues along the way--how much time elapsed? (e.g., deciding to swap client libraries, A/B testing, redis upgrade, shifting load to dedicated instances, implementing proxies/memcached)?
And what was the end user impact (e.g., 50% of users would see timeouts during peak usage for the day, or users using certain functionality would be affected 1% of the time, etc.)
Just trying to get a sense of the level of urgency involved in terms of chasing down all these leads. It seemed pretty methodical, so it's hard to tell if it was a slow-burning persistent nagging issue that you chipped away at over a couple of months, or an all hands on deck sequential process of trying a lot of different things in a relatively short period of time to keep the site afloat.