4 ms·
I once made a huge fuckup. A couple years into my career, I was trying to get my AWS keys configured right locally. I hardcoded them into my .zshrc file. A f
by fishtoaster 2y ago
I once made a huge fuckup.
A couple years into my career, I was trying to get my AWS keys configured right locally. I hardcoded them into my .zshrc file. A few days later on a Sunday, forgetting that I'd done that, I committed and pushed that file to my public dotfiles repo, at which point those keys were instantly and automatically compromised.
After the dust settled, the CTO pulled me into the office and said:
1. So that I know you know: explain to me what you did, why it shouldn't have happened, and how you'll avoid it in the future.
2. This is not your fault - it's ours. These keys were way overpermissioned and our safeguards were inadequate - we'll fix that.
3. As long as it doesn't happen again, we're cool.
Looking back, 10 years later, I think that was exactly the right way to handle it. Address what the individual did, but realize that it's a process issue. If your process only works when 100% of people act perfectly 100% of the time, your process does not work and needs fixing.
- vvanders 2y agoYep, been adjacent enough to a couple large ones through my career to see the details and been up-close to a few that this is the right way to approach it. Did the person know they screwed up? Did they show remorse and a willingness to dive in and sort it out? They likely feel like absolute shit about the whole thing and you don't need to come down on them like a ton of bricks. If that much damage could be done with a single person then you have a gap in your process/culture/etc and that should be addressed from the top. One of the best takes I've seen on this was from a previous manager who when confronted with a similar situation as the article(it was a full DB drop). The person tried to hand in their resignation on the spot, they instead(and I'm paraphrasing here) said: "You're the most qualified person to handle this risk in the future as we've just spent $(insert revenue hit here) training you. Moving forward we want you to own backup/restore and making sure those things work". That person ended up being one of their best engineers and they had fantastic resiliency moving forward. It turns out if you give someone a bit of grace and trust when they realize they screwed up you'll end up with a stronger organization and culture because of it.
- NegativeK 2y agoTo quote a statistician friend: 100% of humans make mistakes. OP's leadership was shit. The org let a junior dev delete shit in prod and then didn't own up to _their_ mistake? Did they later go on to work at a genetics company and blame users for being the subject of password sprays?
- Aurornis 2y ago> they instead(and I'm paraphrasing here) said: "You're the most qualified person to handle this risk in the future as we've just spent $(insert revenue hit here) training you. This is an old quote that has been originally attributed to different people throughout the years. It shows up in a lot of different management books and, more recently, LinkedIn influencer posts. It’s good for lightening the situation and adding some levity, but after hearing it repeated 100 different times from different books, podcasts, and LinkedIn quotes it has really worn on me as somewhat dishonest. It feels clever the first time you hear it, but really the cost of the mistake is a separate issue from the decision to fire someone for it. In real world situations, the decision to let someone go involved a deeper dive into assessing whether the incident was really a one-off mistake, or the culmination of a pattern of careless behavior, failure to learn, or refusal to adopt good practices. I’ve seen situations where the actual dollar amount of the damage was negligible, but the circumstances that caused the accident were so egregiously bad and avoidable that we couldn’t justify allowing the person to continue operating in the role. I wish it was as simple as training people up or having them learn from their mistakes, but some people are so relentlessly careless that it’s better for everyone to just cut losses. However when the investigation shows that the incident really was a one-time mistake from someone with an otherwise strong history of learning and growing, cutting that person for a single accident is a mistake. The important thing to acknowledge is point #3 from the post above: Once you’ve made an expensive mistake, that’s usually your last freebie. The next expensive mistake isn’t very likely to be joked away as another “expensive training”
- vvanders 2y agoI'm fairly certain it occured since the story was first-hand and about 12+ years ago(although they may have lifted it from similar sources). It's not a bad way to diffuse things if it's clear there was an honest mistake Your point on willingness to learn is bang on. If there's no remorse or intentionally negligent then yes that's a different story.
- ay 2y agoSo much this. There is a great book which I think should be on a table of every single person (especially leadership) working in any place which involves humans interacting with machines: https://www.amazon.com/Field-Guide-Understanding-Human-Error/dp/1472439058/ref=mp_s_a_1_1 https://www.amazon.com/Field-Guide-Understanding-Human-Error...
- csa 2y agoIs that a referral link on HN? If so, please remove the referral part. The title is The Field Guide to Understanding ‘Human Error’.
- ay 2y agoNo it is not a referral link. I have searched for the book and then removed anything that looked like it would keep state/extra info. In retrospect, indeed posting a title would have been a simpler option :) thanks!
- kmarc 2y agoBesides the obvious takeaway of the story, to anyone who reads this: use pre-commit hooks to avoid this kind of problems (or something equivalent). With the pre-commit framework, an example hook would be https://github.com/Yelp/detect-secrets https://github.com/Yelp/detect-secrets
- nine_k 2y agoHere's one of may favorite anecdotes / fables on the topic. A young trader joined a financial company. He tried hard to shoe how good and useful he was, and he indeed was, at the rookie level. One day he made a mistake, directly and undeniably attributable to him, and lost $200k due to that mistake. Crushed and depressed, he came to his boss and said: — Sir! I failed so badly. I think I'm not fit for this job. I want to leave the company. But the boss went furious: — How dare you, yes, how dare you ask me to let you go right after we've invested $200k in your professional training?!
- nanoxide 2y agoNever worked with AWS, but besides that it obviously shouldn't happen - is it really that bad? Couldn't the keys be invalidated/regenerated immediately after you realized they were compromised?
- fishtoaster 2y agoOh, they can and were. But bad actors scrape github constantly for access keys. If you commit yours to a repo, some script somewhere will find those keys and use them to spin up EC2 boxes mining bitcoin or use SES to send scam emails within minutes. You can invalidate the keys and scrub your AWS account once you notice the issue - it just depends on how much damage the bad actors are able to do before you do that. In my case, our CTO was messaging me (either Slack or Hipchat - whatever we were using at the time) within an our or two. Iirc they only managed to accrue a few thousand dollars in charges before we got it under control.
- f1shy 2y agoI want a boss like that. Say no more.