3 ms·
This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to reme
by imbriaco 13y ago
This is, sadly, not a great post-mortem. They missed an opportunity for goodwill. I don't feel more confident in their level of understanding or ability to remediate the problems that led to it after having read it. I know they have an excellent engineering and operations staff -- this post-mortem doesn't reinforce that, though.
A few of the things that jumped out at me after one reading:
1. The apology is the next to last sentence. That's burying the lede. I'd like to see that far earlier, in the first two to three sentences.
2. The tone is overly clinical and lacks humanity. I suspect they felt that it made them sound more authoritative and in control, but instead it comes off somewhat robotic.
3. There's a mixture of too little and too much technical detail. It feels like they couldn't decide who the audience was. There were technical tidbits thrown out without any elaboration that lead to more questions than answers.
4. The remediations sound pretty weak. There's no discussion of the human factors like how the recovery process went, how this issue was missed in testing, or what changes if any they think they should make to their incident response process. At the very least I'd expect to see some remediation around their during-outage communication process since it has pretty universally been considered to be poor.
It's not the worst post-mortem I've read, but they missed a few chances to reassure customers.
- kkitay 13y agoI agree with many of your points and appreciate your technical assessment of the actual post-mortem aspect, but your first comment seems particularly nit picky. It's a growing trend that when a company or person fucks up, we expect a big, grandiose, sobbing apology (and when they don't, we blow a gasket - a la Snapchat). Now, I'm not saying that I don't expect companies to be forthright and take ownership of their mistakes, as well as apologize for them, but I can't help but feeling that expecting Dropbox and others to get on their knees and kiss their users' toes when something happens is a little melodramatic. On the one hand, yes, they made a mistake - on the other, we all know that technology is flawed, and these things happen, albeit rarely. TL;DR: Let's not make a drama out of it.
- imbriaco 13y agoI'm admittedly being nit-picky because I feel very strongly about the importance of outage communication. Good communication both during and after an incident can make a tremendous amount of difference in how you are perceived. They decided that it was worth apologizing for near the end of the post. All I'm suggesting is that moving that up near the top and acknowledging up front that they let customers down would have improved the outcome. They don't need to be over the top about it, just don't bury it at the end of the post.
- kkitay 13y agoFair enough. Your comment was more of a spark of a sentiment I've been carrying around for a little while. I can't agree enough that proper outage communication is important.
- kordless 13y agoAs I've said before, one blog post does not represent that team when it's written by someone tasked with the job of communicating with a wide variety of customers. My mom could give two hoots about details. She wants to know why her 'spinny drobox thing' keep spinning and should she upgrade or something. I deliver that news to her. This blog post delivers it to people who don't understand as well as most of us but better than my mom. What would be an AWSOME idea is if Dropbox did a meetup to go through the gory details for us nerds. Now that would rock. Kudos the the Dropbox team for working through the weekend fixing stuff. I spent the better part of the weekend nursing a barely 2 year old dying Apple 27" Cinema Display back to life by disassembling it several times. Kept thinking to myself that I sure as hell was glad it wasn't me over at Dropbox HQ working on doing recovery instead. Edit: I agree with your plea for emotion in the post. It could ease things a bit.
- imbriaco 13y agoThat's just it, though. This is the public face of the team that responded to that outage. It absolutely represents them. Now, whether it's a fair depiction or not is definitely a valid question. Having written more than my fair share of these, I definitely understand the difficulty involved in choosing your audience and writing to them. That's a big part of the problem here: The audience is not clear. It bounces between technical detail like MySQL recovery process, but it doesn't go deep enough to be satisfying for a really technical audience while being too detailed for a non-technical one. I have nothing but admiration for their team and the service they've built, but this post-mortem misses the mark.
- kordless 13y ago> the audience is not clear. Bingo. We need nerd updates. BTW, we deserve this because enough of use use Dropbox for quite important things coding-wise.
- _pmf_ 13y ago> There's no discussion of the human factors like how the recovery process went, how this issue was missed in testing, or what changes if any they think they should make to their incident response process. Not every company is into that whiney startup blood and tears thing. Those "we() worked non-stop for the last 72 hours" often sound a bit desperate. () And by "we", the PR people usually mean the engineers.
- imbriaco 13y agoThat's not at all what I was getting at. It's not about patting yourself on the back or trying to make the team look like heroes. I was more wondering how the mechanics of their incident response processes were managed and whether they planned to make any changes as a result of the review of this incident. Technical remediations are all well and good, but organizational, cultural, and even procedural changes are often even more impactful after events like this. For example: Were they happy with the pace of communication during the outage? Do they think customers were updated frequently enough, too frequently, etc. Any changes planned? How did they handle incident fatigue? Did they have to go to shifts to manage the recovery? Did they already have this planned or was it done on the fly? Do they plan to build any procedures to handle similar long-running events in the future?
- chris_wot 13y agoI don't see the need for any of that. What does it really matter they had "incident fatigue"? I don't really care about their internal comms or escalation procedures. If I was a customer, I'd want to know what they are doing to mitigate a similar incident (which they answered), and an apology. If I wanted a credit, or SLAs weren't met, then I'd talk directly to an account manager.
- wpietri 13y agoThe thing I always look for in post-mortems is an understanding of the failure of human systems. The technical failures are interesting, but it is the human systems that produced the technical failures. And will keep on producing other failures unless changed. I hasten to add that I'm not looking for finger-pointing or blame. In retrospectives, I think it's always best to assume that individuals did the best with what they had. [1] But I think it'd be great if Dropbox asked themselves things like "How did we miss this bug?" and "How could have we discovered this recovery issue before it was on the critical path for a public outage?" Questions like that help you solve not just this bug, but all the related latent bugs that you got the same way you got the one that just blew up. [1] A lesson I learned from Norm Kerth: http://www.retrospectives.com/pages/retroPrimeDirective.html http://www.retrospectives.com/pages/retroPrimeDirective.html
- mkagenius 13y ago"How did we miss this bug?" and "How could have we discovered this recovery issue before it was on the critical path for a public outage?" I am pretty sure they would have done that - just that they did not include in the post mortem.
- chris_wot 13y agoThe first was answered in the postmortem. The second is something done in time - either it's hard to answer in detail without revealing confidential information, or they are working towards it in the medium term.
- RyanGWU82 13y agoThis post was just an incident review for a technology audience. Dropbox posted a separate apology to their users on their main blog: https://blog.dropbox.com/2014/01/back-up-and-running/ https://blog.dropbox.com/2014/01/back-up-and-running/ . The tone and detail seem totally appropriate since it ran concurrently with the other post.
- imbriaco 13y agoAh, that's interesting. I wasn't aware of this post at all, thanks for pointing it out.
- colechristensen 13y agoThis is a silly fetishizing of 'post mortum' I will think no less of any company with solid technology that experiences a failure and puts an honest effort in communicating a post mortum explanation which is exactly what happened here. I will, though, lose some respect for people who quibble about perceived faux pas of the explanation because it's losing sight of what is actually important.
- VMG 13y agois this satire?
- WestCoastJustin 13y agoGoogle provided a great Incident Report / Postmortem when they had their API infrastructure outage back in May 2013. I created a screencast about how their template should be used as a model for the rest of us to follow. You can watch the screencast @ http://sysadmincasts.com/episodes/20-how-to-write-an-incident-report-postmortem http://sysadmincasts.com/episodes/20-how-to-write-an-inciden...