5 ms·
I had a jr dev connect and typed 'flushall' because he thought it would refresh the dataset to disk. thankfully it was on a staging env, I think he's at google
by scrame 3y ago
I had a jr dev connect and typed 'flushall' because he thought it would refresh the dataset to disk.
thankfully it was on a staging env, I think he's at google now.
- spacephysics 3y agoOne time during my internship years ago I took down a production server because of a command I ran on it that I didn’t fully understand. Since then I treat any prod server terminal like I’m entering launch codes for a middle system. Anything outside of ls or cd I’m very careful, read the command a couple times before executing, etc.
- returningfory2 3y agoIn my opinion you weren't at fault here. Production systems should be designed so that one person can't inadvertently destroy things.
- RyanHamilton 3y agoThis is the way.
- ljm 3y agoIn almost every place I've worked at, the most difficult thing has been getting people out of ad-hoc JFDI style development and debugging, where everything in production is fair game, and into a process where you avoid touching production as much as humanly possible. Takes a lot of effort to stop people opening up a shell in prod or grabbing a prod DB dump or even just connecting to the prod datastore directly from their local env.
- deleted 3y ago[deleted]
- tetha 3y agoFor critical and overall... fiddly things, we've grown into a culture of writing down reviewable plans and possibly executing these plans in pairs. We tend to go ahead and either use a runbook, or whatever experience we might have, to setup a pretty detailed plan of what to run on which systems with which purpose. You can then throw these plans at someone else to review. Sure, it takes an hour or two more to setup a solid plan and waiting for a review takes time as well. But this has turned into a great tool to build up experience in weird parts of the infrastructure.
- scrame 3y agoI get that, but it also can turn into wiki checklists of things that could be automated. one of my frustrations of bigco software is people taking basically maintenance roles where the computer tells them what to do, because the lava flow legacy code base is too scary to touch. however, you can automate your daily clean up tasks. it's certainly shellacking more mud on the ball, but if you're not going to even try scripting your repetitive tasks, then i don't know why you're a programmer.
- tetha 3y agoOur running gag is: Once such a runbook has been sufficiently refined and clarified to the point of being really comprehensive and easy to follow.... someone turns it into a jenkins job and we don't need it anymore.
- js2 3y agoYou can use rename-command to help avoid these kinds of mistakes: # To disable: rename-command FLUSHALL "" # To rename: rename-command FLUSHALL DANGER_WILL_ROBINSON_FLUSH_ALL
- cik 3y agoInstead, use the redis acl commands to modify permissions you don't want executed by a user (i.e flushall, flushdb) as opposed to renaming commands.
- signatureMove 3y agoif only my immaculate record of never "rm -rf"ing myself or prod dbs resulted in me working at google...
- mtlynch 3y agoIt sounds like the subtext is that this dev was incompetent, but if all they did is mess up a staging environment, it sounds like things were working as intended. If a junior dev can cause catastrophic harm from one wrong command, it's the org's fault for not having safeguards in place, not the dev's fault for an (understandable) error.
- scrame 3y agoyeah. we built many many moats of protections, but this guy was... optimistic, I guess? by and large our redis stuff was ephemeral, but we had a particular key that was a domain table that needed to be loaded separately, and that caused some problems. incompetent is maybe a bit harsh, but i did say he was junior, and junior devs make mistakes, and this guy was well meaning and messed up. you don't get from junior to senior or principal or staff without some mistakes, and it's the responsibility of the more senior devs to not have them in a position where their mistakes are catastrophic.
- yawaramin 3y agoHe learned a very important lesson. That's one guy you can pretty much guarantee (if he has any brains at all) will be very careful about doing anything on a live production system in the future. In this case Google probably got a good deal.
- stevekemp 3y agoReminds me of the time I ran "killall" on SunOS, which didn't kill a process by name as it did under Linux, instead it killed all processes. That's the kind of mistake you only make once!
- squeaky-clean 3y agoI've had someone do this in production. Even worse, it turns out when each microservice needed a redis instance, sysops was just expanding the main redis instance and pointing the service at it instead of giving each microservice their own instance.