7 ms·
I deleted data from production
- blef 5y agoI share 3 stories about my previous experience when I deleted not on purpose data in my Hadoop clusters or when we ran terraform destroy over our GCP projects. This is a personal post about what I learnt and some takeaways. Did you ever feel the same way?
- sojasun 5y agoThank you for sharing your experience !
- rnotaro 5y ago> Create a good wheel environment, ask for help and *do hide stuff from colleagues* Quoi? Do *not* hide stuff from colleagues. That's probably what ended up causing `terraform destroy`. Do not hide stuff from your coworkers. It's toxic in a workplace. Hopefully it's a typo.
- masswerk 5y agoI'm pretty sure it's meant to include "not" with regard to the ethos of the entire article. (Ought to be fixed though.)
- betaby 5y agoIAC promise was ease of re-creation. [My anecdotal] Reality shows that's not really the case. IAC automation helps to create 'easy parts' and to delete _everything_. Creation of hard pars, the ones which involve state is still a problem. After many ears of using cfengine/chef/puppet/ansible/terraform I can say those tools don't help much with solving infra complexity. Early arguments were that 'nobody understand those shell scrips' or 'know manually edited configs'. Now the same shell script live inside container 'sidecars' and do 'pip install' from them and instead of manually altered configs we have k8s annotations with unpredictable 'far reaching' inter-references.
- jjnoakes 5y agoThe hard stateful parts should be your data layer (replicated db, etc) and everything else should be cached stateless data. IAC for the stateful layer is tough but contained. IAC for the rest should be fairly recreatable quite quickly.
- KronisLV 5y ago> Now the same shell script live inside container 'sidecars' and do 'pip install' from them and instead of manually altered configs we have k8s annotations with unpredictable 'far reaching' inter-references. This sounds really wrong - installation of all of your dependencies should be run as a part of building the container for a particular version of your application, say, as a directive within the Dockerfile. That should also be done before any unit tests and integration tests are run against that container, well before it is deployed to any environment. The end result of doing that properly is a container that will work correctly until the end of time (or at least any vulnerabilities become apparent and a new version needs to be built, or for other reasons, such as bug fixes or added functionality). Furthermore, all of your IAC should be treated just like regular code, which means full descriptions and comments for any non-trivial parts, or even explaining why the trivial parts are there to begin with (the greater circumstances, links to change requests etc. that the actual code doesn't otherwise contain). If you can't look at any single piece of your code and understand both what it does as well as why it's attempting to do it, then clearly there is still more work to be done. Whether that's refactoring or documentation, however, depends on the circumstances. IAC doesn't absolve you from the need of having documentation. It simply allows you to lower the degree to which PEBKAC is a problem in manual deployments and allows you to document the actual code that does things, rather than creating .docx descriptions that no one will read.
- xtracto 5y agoA friend of mine once ran the Ansible equivalent of terraform destroy in our aws prod env in the middle of the day (of a fintech company). I could see how several if our prod machines got terminated. Fortunately the ones that mattered the most, had termination protection. But it was a scary moment.
- psim1 5y agoThis is why he is Data Engineering Coach and not actually responsible for production systems now. (apparently) Everyone loves reading tech horror stories and peeling off the take-away lessons. My lesson would be, don't hire this guy and don't use him as a coach. He's careless! "But this kind of thing could happen to anyone" - sure, anyone who is irresponsible. Am I ranting? Let me continue. I work on a team now with someone who seems to consistently forget to save router configurations and another person who yawns through meetings because he doesn't sleep. These guys make careless mistakes and we "learn" from them. Except we don't learn the most important lesson: Irresponsible people need to be off the team. They make more work for everyone else and make the team look bad. But I guess the tech horror stories that amuse everyone on the forum or at the pub are better than saying, "I am responsible and do good work."
- whatever1 5y agoNone of the code for consumer production is verified formally. So please spare us the bs that you make no mistakes. You are just lucky. Be humble because the complexity of modern systems is insane, there is no way you have all cases covered, if you had you would had a formal proof. We all are just doing our best to cover most of the edge cases. That is why we need to keep learning from other peoples mistakes.
- sdhfjg 5y agoRead the article. We're not talking about formal verification. >It's Sunday morning and I just discovered that I've lost 3To of data and that all data pipelines have stop working because on Friday I ran for no reason hdfs dfs -rm /data This is profound incompetence.
- a012 5y agoThe stories scream incompetent people being incompetent.
- tomcooks 5y agoThis is precisely what one thins before doing an honest mistake, then they learn they're less perfect than expected.
- kelnos 5y agoI agree with that in principle, but I think the things OP writes about are egregiously bad. Both in the sense of the employees in question being incredibly careless, and the company being negligent in developing process around having this sort of "god mode" access. I've seen my share of honest mistakes (and committed some of them), but... damn.
- chromatin 5y ago> Never blame the responsible, I prefer to think that if the mistake happened it's because the team or the company let the issue happens — be also careful when you joke about it afterwards Although certainly people still do bad things, dumb things, and careless things -- this (the middle part -- team/company controls and culture) is a wise management approach, although best applied as a principle ahead-of-time.
- winternett 5y agoI lost a PM job once because a new-hire junior admin deleted a symlink and took a major site down... The CEO who fired me was upset that I couldn't restore the massive database and run a diff on it versus a local backup I made out of being overly cautious the day prior. I was supposed to just be a PM, not even supposed to be a Dev... Served me right for acting like I might have been able to fix the issue. No more volunteering outside of my role... Possibly also why job descriptions include so much out of normal scope "responsibility" now. The company had a bad habit of underpaying roles, and regularly hired inexperienced admins to do key PROD work for clients. I was also blacklisted for rehire from the company and took the fall I guess when the issue was explained to the clients. Oh well.. Spilled milk... The company kept failing after I left and lost all it's major contracts weeks after I got cut anyway. I got a decent severance payout and still won a 25% raise on my next gig... Moved forward without scratches. Lessons - Back everything up well, no matter how much time it takes, before touching PROD or even before others touch PROD. Don't work for places that are regularly dedicated to underpaying talent to do mission critical work. When a door closes, another opens, learn from the past, but don't carry it forward, and don't fight ignorant and tech-blind arrogance in leadership, it only hinders your career growth if you crusade against bad leadership. Karma does the real work for you eventually if you remain calm & composed. Quickly identify if tech talent and problem solving stops at your level, and if so, protect yourself from being a scapegoat for leadership failures. Don't be the one that screws things up, be the one that both anticipates and quickly and fixes all varieties of screw-ups from everyone else.
- thaumasiotes 5y ago> I lost a PM job once because a new-hire junior admin deleted a symlink and took a major site down... The CEO who fired me was upset that I couldn't restore the massive database and run a diff on it versus a local backup I made out of being overly cautious the day prior. But... the database wasn't gone? If you delete a symlink, you can just put it back.
- winternett 5y agoThe (devops) admin deleted the link to an assets folder for all the file uploads for a pretty large web site and didn't even know how to explain what he did. We all didn't know it was a symlink that was deleted until later after the failure, but the disconnection caused the entire site's db to go corrupt... I restored the link but that didn't solve the resulting corruption, so a DB recovery/diff was in process from my desktop backup when they turned to blame on me...The DB was really FUBARED... Actually quite glad I didn't have to finish the fix on my own after all that drama TBH. :/
- eftychis 5y agoIf you are working later at night over the weekend something bad is going to happen (except if you work on a Wednesdays off Sundays working days schedule). If you mix production with dev something bad is going to happen. If you expect no mistakes to happen, well, something bad is going to happen. Authorization and prompts are there for a reason. @comments discussing about dismissing OP -- doesn't matter: you don't have the data anymore. Production data should be indestructible practically or it doesn't matter. Sure your client/user is going to be OK if you fired your junior dev Bob... Just comment how many times that has worked.
- deathanatos 5y agoI've had someone do, in essence, a `terraform destroy`. I don't think we ever learned if it was that, or some other command … just that it was indeed a terraform removal. The plan was not read, of course. Instant incident. More recently, had a coworker want to run terraform plan, couldn't, because the state file is access controlled, and so he just reset the configuration locally to a blank local state, plan (which wasn't read, again), and then subsequently created a lot of duplicate resources. (Since TF was essentially starting from a blank slate.) We had an admin interface that used a generic UI; every database table was just a generic CRUD "it's a table, here's the columns & values" to it. The mobile app was an OAuth app, just like any third-party app was. One day, someone on the QA team (why did QA have access to prod, you ask? Good question…) deleted the mobile app's OAuth client from the database. "Why did you do this?" :shrug: "Why do we even have this lever, Kronk?" I suggested at the time we should have kept the delete button, but special cased it to remove the admin privs of anyone who clicked it, since clearly they failed the test… I can't tell why the copy in the article didn't work? You should be able to move /usr to another partition, no? (I'm not sure I'd want to try that on a live OS … and … if this is the cloud, usually growing the underlying EBS or cloud disk or whatever is pretty easy to do? But I also usually try to keep any "it could grow" data on a separate partition to spare / the pain…) Though I think I'd bind-mount it back in place, not symlink it.
- anyfoo 5y ago> I can't tell why the copy in the article didn't work? You should be able to move /usr to another partition, no? The way it reads, they didn't know how POSIX permissions work, and maybe still don't. The article says that sudo needs to be owned by uid 0, but does not mention the critical suid bit. Maybe they just recursively copied /usr without preserving permissions? Also, the mention of sudo makes it pretty clear that they did not do this in an interactive root shell (obtainable with sudo -s/-i or su if there is a root pw). If you do this kind of shenanigans, you better have multiple open root shells in front of you, and some statically linked binaries (busybox?) to recover from the inevitable chaos of missing binaries, shared libraries, and dynamic loaders that you'll transiently have.[1] It's not impossible to perform, but requires careful planning and a good understanding. [1] Though if /usr was indeed copied without preserving permissions, recovering would have been very tedious.
- mustardo 5y agoMy favourite was when a "helpfull" sysadmin deleted all the ActiveMQ data(base) files for a prod system (containing unprocessed queue items) I assume they got a low disk space alert and started deleting big files with no notice, recreating the messages from logs was fun :{
- kelnos 5y agoI'm really torn on this stuff. On one hand, the big failure is a company process one: people should not have the permissions to just run "terraform destroy" and have it actually destroy a production environment (ditto for the HDFS /data deletion). Or at the very least, there should be a strong culture of never getting into the position of being able to do that without someone looking over your shoulder, double-checking every command before you run it. But... some of this just feels like carelessness. I've certainly made my share of mistakes, but these just feel egregiously bad. This takeaway in particular is just all wrong: > Measure the risk when you give all the permissions to one developer, one data engineer or one SRE — it means similar stories could happen There is no need to measure that risk, because no one should have permissions to take down your production infrastructure so easily. This just seems like table stakes for running or working at a company that operates this kind of infrastructure. I think it's three things: first, abide by the principle of least privilege, and make company-destroying permissions hard to come by; second, put safe (web, command-line, whatever) interfaces in front of common tasks that need to be done that could turn into accidental downtime; and third, when it's required that you do things outside of the safe interfaces, drill it into people that you never do them without a copilot who can check over your work, in real-time, before you do anything (and if anyone develops a reputation for being a cowboy... seriously, fire them).
- noisy_boy 5y agoThere are three examples in the article: removing /usr (!!), running "hdfs dfs -rm /data" and "terraform destroy". Even with basic knowledge, these are clearly destructive operations - and to top it all, being run in Production. What strikes me as odd is that in all the cases, there is no change control. And by that I don't necessarily mean a full-on ServiceNow-type change process; I mean there was no "review". Even if you are a small firm, it is not difficult to ensure that you ask atleast one colleague to review any destructive step in Production prior to executing. Ignoring to do so is inviting disaster.
- thatwasunusual 5y ago> It's ok to do mistakes. Not _these_ mistakes, though. The author describes incompetence, nothing else.
- bob1029 5y agoInstead of playing an elaborate finger pointing game, we decided to develop internal tools which put our most dangerous operational activities on rails. It's amazing how much impact a small console application can have on preventing mistakes. All you really need to do is wrap the danger with "are you sure?" and "enter the exact commit hash..." kinds of prompts. For the super dangerous stuff you can bake in some centralization and approval loops.
- wiz21c 5y agoTalkinf of personal stories, I'd be very happy to have a tool that enforces some verification process over SQL operation ran on production database (you know, to fix that horrible data quality issue on 2 rows that screws the whole business logic in the upper layer that would take a month to fix). Basically a tool that allow someone to run some SQL on production but, before doing it, request his credentials, then run the sql command in some "--dry-run" mode and if you modify more than 10 rows, then it stops everything. And if it goes on, record the SQL that was run somewhere so we can audit...
- idoubtit 5y agoAs a first step toward your wish, I suggest using mycli or pgcli on servers instead of the default mysql or pgsql client. By default, it asks for confirmation whenever you type a destructive command. It does have a log, just like the official CLI clients. It also provides completion and syntax highlighting. https://github.com/dbcli/ https://github.com/dbcli/
- bullen 5y agoMy database distributes data in real-time globally: http://root.rupy.se http://root.rupy.se That way it's almost impossible to have anything catastrophic happen unless the whole planet goes boom!
- KronisLV 5y agoThat's an interesting project, however you might want to consider slightly altering the CSS for the homepage: http://rupy.se/ http://rupy.se/ Currently the main font: font-size: 0.6em; ends up being around 9.6px on Firefox with default settings in regards to font size, which makes it pretty much unreadable on a 1920x1080 monitor (21.5"). Example screenshot: https://imgur.com/IsPvHqE https://imgur.com/IsPvHqE
- ricc 5y ago> Create a good wheel environment, ask for help and do hide stuff from colleagues (emphasis mine) Was this a typo? If it is, it's kind of funny that an article about making mistakes in production also made this tiny but critical mistake.
- blef 5y agoYes, oops.
- meowfly 5y agoI have sympathy for the author and find a lot of the criticism a bit harsh. On my teams, I tend to be the person put in the role of doing DevOps-ish work because it's something I'm familiar with. I honestly find it incredibly stress inducing. It seems to me there is an asymmetry in the risk of work between team members that never gets captured by management. There are degrees of difference with security and data concerns in regards to frontend versus backend versus devops tasks. There is some work that can carry the risk of permanent and irreversible data loss, yet everywhere I've worked, it is apportioned as a regular tasks. Failures like Ops are first and foremost an organizational issue.
- blef 5y agoThanks.
- blef 5y agoHey, OP here. I wasn't ready to get so much negativity from my post. OK the title was a bit clickbait but I just shared honest feedback because I was sure that a lot of people could recognize from my stories. I never said I was an expert in the post, I never said I consider myself better than everyone, and moreover I'm 100% sure I can bet I've not impacted the life of one person that commented my post. Is it my fault when I just graduated that I had been given the responsibility of building an Hadoop cluster for a +$100m revenue company? No, I just tried to do my best. I fucked up things and I fixed it. Period. I'd love to work with all people telling me I was incompetent, but unfortunately you weren't here to bring me the light in 2014. 7 years after I can sleep at night and live with my mistakes. So, yeah, sorry for that. I'm closing this, love.
- otagekki 5y agoIt reminds me that back in July 2019. I was using the Windows file explorer to investigate an issue. When clicking on a folder to open it, my wrist slightly twitched while clicking on the folder, which led to the accidental move of a 650 GB operational data folder into a code repository folder. For some reason the move into the repository folder was instantaneous but the opposite wasn't. All data import attempts errored during the 14 hours required for the data to be moved back to its original place. I emailed and voiced my apologies to the team present both in India and in France. While I was not fired on the spot, it might have indirectly led to the disappearance of our devops (devoops?) team in Paris. Fortunately enough I guess I was the last team member to be sent away. For my defense, our dev team had a intervention scope so big that they were essentially super-men once they have been granted access. One the rights granted to us was full write-access to all incoming data folders.