6 ms·
sales never takes the blame. If anyone is fired it will be scapegoats in engineering once they have busted their ass to restore their reward will be the door
by syshum 4y ago
sales never takes the blame. If anyone is fired it will be scapegoats in engineering once they have busted their ass to restore their reward will be the door
- systemvoltage 4y agoThis is an engineering problem. They should own it and improve things, make sure it doesn't happen again. Also, GP's quote > Engineering mistakes happen. I don't like this statement because it offers consolation at the expense of unintentional normalization.
- buscoquadnary 4y agoAnd coders that say all code has bugs are just defeatists that are trying to make excuses for being lazy. Sometimes manure will always hit the fan. Being robust means being able to handle that.
- systemvoltage 4y agoA culture where mistakes are taken too seriously or too lightly leads to problems. Also it depends on what stage of the product cycle (Innovation/Rapid Development vs. Robustness/Quality). I'd argue that Atlassian products should err towards robustness and high quality. Not trying to break any new ground.
- jacksnipe 4y agoI think this is obviously incorrect. Human error is probabilistic, and the probability of making an error cannot be zero. On the flip side, it’s infeasible to use only provably correct systems; not lazy, but literally not a practical option due to compute costs, developer time, what formal techniques can even be applied to the problem at hand, etc…
- tinco 4y agoIt's not about making the probability of an error zero, it's about making sure you can recover from every type of error effectively. They've underbudgeted for engineering, and they're feeling that now.
- josephg 4y agoSure human failure is probabilistic. But you can design around that by stacking reliability-enhancing approaches together. Let’s say there’s a 10% chance of any given feature being broken. Write a test, (which has another 10% chance of being broken) and now it’s only broken if the test and the code are broken, and broken in the same way. So we’re down to <1% chance of failure. In my experience most bugs that make it past testing do so because you forgot a test. Then add a backup / redundancy system. That has a 10% chance of failure, but if you test it regularly then the backup / restore process only has a 1% chance of failure. Now we have a system that’s pretty reliable in practice, made out of pieces which are only 90% reliable. And no need for PhD level formal methods. Just do the obvious robustness steps: Write unit tests. Run them with every commit. Have a backup system. Test it. Have redundant servers. Do stages deployments. Monitor your servers and have an on-call roster. Then when everything is working well, add a chaos monkey to increase the failure rate of all of these parts so your team & software gets practice dealing with problems. The fact that this bug slipped past all of their reliability engineering - past code review and testing into production and in a way they can’t recover - that smells of sloppy work.
- jimbokun 4y agoThey had backup restore process. The trouble was the restore would set back everyone’s data to that point in time, whereas only some customers data was impacted.
- throwawayboise 4y agoI wonder if in retrospect that would have been better. If they had rolled back to a snapshot 30 minutes after they realized they had a problem, everyone loses 30 minutes of updates (and maybe transaction logs can be copied before the rollback and then replayed to reduce that to even less). Everyone experiences a little bit of pain instead of some customers being down for a week plus. Easy to speculate about from the cheap seats though.
- calsy 4y agoIt's the exact opposite, any coder who blindly believes that a piece of software is flawless is kidding themselves. It's delusional to think software can be flawless in the real world when it's used by an untold amount of people, on all manner of devices, possibly running different OS's with different versions on networks that can be configured all sorts of ways. Thats not to mention all the dependencies involved in creating high level software, from the third party libraries to external services like cloud storage. You anticipate there will be problems and make sure there are processes in place to manage them when they inevitably occur. Thats the exact opposite of laziness.
- lelanthran 4y ago> Sometimes manure will always hit the fan. Being robust means being able to handle that. You're never going to get perfect error handling in any non-trivial system. Being robust means that you plan for particular states (like "deleting the production data"). That doesn't mean that your plan is any good, or that your plan will fix the problem, only that you have a sequence of steps developed in advance of the problem. Sometimes the state in question is considered too unlikely[1] to ever occur, so is ignored with the caveat "too unlikely", such as planning for the case when the company files for bankruptcy and all software needs to be sped up by a factor of two in order to halve computing costs. Not all possible future states need to be accommodated for in the tech stack - that doesn't mean stack is not "robust". [1] Or if likely, is such a large problem that all the other problems are irrelevant.
- nix23 4y agoEver heard of Space Shuttle Challenger? You cant own it if your management is against it.
- syshum 4y agoThe deletion of customer data was engineering mistake, that is not what I was talking about The Negative fall out was not due to the deletion of customer data, as the Story and multiple customers have stated the negative fall out was the SILENCE / lack of communications, which is Sales / Customer Service not engineering As the comment I was replying to noted while engineering was trying to recover from what might possibly be the biggest outage in the history of the company Sales was partying and not handling customer communications That (the failure to communicate with customers) should be a resume generating event of all leadership customer service / sales. It will not be because sales will simply redirect their failure on to engineering in the exact same manner you just have
- systemvoltage 4y agoOk I agree with these failures, but don't you think that its a PR people problem? Perhaps executives and upper management? Sales people are just doing what they are supposed to do. Sell Atlassian products.
- rrook 4y agoYou and the person you're replying to are using the word "Sales" differently. GP is using it as "Sales Representative", a la Jim Halpert, whereas you're using it as "Outbound Sales", like Glengarry Glen Ross.
- fphhotchips 4y agoReformed salesperson here: bullshit. Atlassian famously eschews the exact sales teams whose job it would be to manage direct customer comms in an outage like this one, and to be the lightning rod for the understandable customer frustration. In the past, I've been the guy that gets the angry text message from the customer and has to carefully paper over the gaps in communication from higher ups. It's not fun being the neck that gets choked. The complete lack of meaningful communication for so long indicates to me that Atlassian doesn't have a meaningful feedback mechanism from the field back up to the executive suite - the exact feedback mechanism that Sales and Sales Engineering teams fill in most SaaS orgs. Customer Success should fill that role, but IME don't have the same incentives, pressure or influence as sales teams watching half their yearly comp go down the pipes.
- zelphirkalt 4y agoThey also seem to eschew a proper UX design team, quality assurance and even proper engineering, judging how "well" their software works.
- Karrot_Kream 4y agoWhy should sales take the blame when it's engineering's problem? If Sales promised a feature to a customer that was infeasible _then_ it should be Sales problem, but engineers made a mistake so only engineers can clean up the mistake.
- theteapot 4y agoWhat? Because sales interacts with customers and the customer doesn't give a F*$% who's fault it is or even know engineers exist.
- Karrot_Kream 4y agoThis depends on your org structure. Sales does not focus on ongoing relationships at our company, instead we have dedicated account managers who handle relationships with large customers and have a general process to publish status updates during an outage. The outage status update process is guided by an internal outage management team and is a big deal; our C-suite and VPs immediately action on gaps managing status updates. We promise SLAs so managing expectations is key to hemorrhaging money from outages. Without knowing Atlassian's org structure all I, an outsider, know is that engineering had an issue and remediation status communication has been lacking. If anyone should take "blame" it should be the bad engineering practices (usually a lack of funding or inability for upper management to prioritize engineering risk) which led to an outage of this magnitude and the status update mechanisms in place. Even then, I question how much better status communication would even get Atlassian given the sheer length of this outage, which is why I think the blame should ultimately land on engineering practices. Hopefully Atlassian learns from this but I'm afraid the damage is done. A SAAS system which takes this long to heal is unacceptable in 2022.
- syshum 4y agoIn many orgs Sales and customer service is the same team, under the same leadership. However in the context of this conversation we are specifically talking about the "Global Head of Customer Success" if you took the time to read the comment I was reply to. Your reaction (and a few other clues) leads me to believe you are in sales... I am i Right???? In the contest of this outage, the "Global Head of Customer Success" should absolutely be taking the blame for failing to communicate with their customers. You can attempt to deflect the blame purely onto engineering side of things, as people in sales often to, but you are objectively wrong. I have suffered long outages with vendors before, the key for me was always COMMUNICATION. They should be communicating with every customer impacted, they should be telling them on going progress, where in the queue they are, etc etc etc. Trust me as someone that make purchasing choices and recommendations the length of this outage is not the issue, the lack of communications is. As someone that makes purchasing choices and recommendations I do not blame the engineering, I blame sales or "Customer Success" or what ever PR name you want to further deflect sales to be....