4 ms·
> The engineer fixing the Azure Table storage performance issue believed that because the change had already been flighted on a portion of the production infras
by sandis 12y ago
> The engineer fixing the Azure Table storage performance issue believed that because the change had already been flighted on a portion of the production infrastructure for several weeks, enabling this across the infrastructure was low risk.
Ugh, I wouldn't want to be that guy (even if there would be no direct repercussions).
That said, and as others have highlighted - kudos on the writeup and openness.
- cheez 12y agoForever known as the guy who knocked Azure offline.
- click170 12y agoProbably less than you think. At least among his peers. These kinds of "almost took X offline" happen All The Time, its just that most of the time they get caught before it gets too far. Its inevitable that a few will squeak through the nets. Mistakes can and will happen anywhere we allow them to. If you want to prevent mistakes, write tools to help reduce the "attack surface" (areas where mistakes can be made). Eg Don't want someone to be able to do "sudo reboot" accidentally? Alias reboot to something else. It won't stop hackers but it might help fight fat fingers.
- epochwolf 12y agoI've halted production in 14 manufacturing plants before. Ran a query against the wrong server and locked up the entire ERP system for 30 minutes. Accidents happen.
- angersock 12y agoIt's kind of shitty--it kinda seems like it had passed initial testing, and then got rushed out to full-scale production because it also fixed other customer issues (as mentioned in the writeup). That's the thing about these kinds of bugs...they are, by definition, tricky enough to have passed testing unseen.
- jschmitz28 12y agoThe sentence you quoted seems ambiguous to me, since (from what I understood reading the article) there are two separate storage mechanisms using the new feature. The two possible beliefs the engineer may have had are: 1. Since we tested the change on a subset of A for a few weeks, we can assume it will work for all of A. 2. Since we tested the change on a subset of A for a few weeks, we can assume it will work for all of A and all of B. #1 seems reasonable, but #2 is what needed to hold true in order for there to be no problems, since the change was actually enabled for all of A and B. But was the engineer actually advocating to enable the change in B, or was that an accident during the manual deployment?
- nchelluri 12y agoI don't agree with your possible beliefs. What about: 3. Since we tested this change on a subset of A and a subset of B, we can assume it will work for all of A and all of B.