5 ms·
Salesforce outage: Global DNS downfall started by an engineer trying a quick fix
- cratermoon 5y ago10:1 some manager or product/salesperson told the engineer to rush it.
- justin_oaks 5y agoAnd that person is probably is the first to blame the engineer too.
- deleted 5y ago[deleted]
- Mandatum 5y ago> "We're not blaming one employee," said Chief Availability Officer Darryn Dieken > "For whatever reason that we don't understand, the employee decided to do a global deployment," Dieken went on. The usual staggered approach was therefore bypassed. > And the engineer who sidestepped Salesforce's carefully crafted policies and took down the platform? "We have taken action with that particular employee," said Dieken. Holy contradiction, Batman!
- viraptor 5y agoI don't know what Dieken had in mind exactly, but one interpretation that's not a contradiction could be: "The employee did a clearly stupid thing, we blame them for that. We don't blame them for the outage which we could've contained at multiple earlier steps." Again - only playing devil's advocate. We'd need to know much more about processes and what actually happened for a better explanation.
- higeorge13 5y agoTo be honest the whole article reads as a straightforward blame-the-engineer speech. They refer to him multiple times and eventually we learn that they have taken action (everybody can guess what this means).
- oconnor663 5y agoIt might also be that they're taking a CYA tone in public but a healthier tone internally.
- Aqueous 5y agoreally dislike the doublespeak here. speaks very poorly of this man’s leadership. either blame the engineer or don’t, but don’t try to have it both ways
- justin_oaks 5y agoThis looks pretty bad on Salesforce's engineering culture. 1. They're still using manual processes where automation should be used. 2. They're using insufficiently robust scripts (Forgivable to a degree. Bugs happen) 3. They blame the individual rather than the process which allowed the individual to make this mistake. 4. They have their status page on the same infrastructure that the status page is reporting on.
- supergirl 5y agosure, blame it on 1 person. don’t blame the company with such a messed up infra that 1 person can accidentally bring it all down
- chtitux 5y agoTo be honest, DNS is probably the only piece that can easily take the entire infra down (including the staging one). It's so easy to rely on DNS names everywhere...