3 ms·
If that administrator is reading this, chin up ... it happens to the best of us.
by _spoonman 11y ago
If that administrator is reading this, chin up ... it happens to the best of us.
- bpchaps 11y agoIndeed. My own examples: 1. Accidentally restarted a bank's FX platform when troubleshooting a failed cron. I copy/pasted/returned, "vi source ~/.bashrc ; ~/scripts/restart_env.ksh". Nobody noticed, but I still had to make a dozen cold sweat phone calls. 2. Same place.. through their GUI, effectively ran a "select * from table1,table2,table3,table4,etc". The entire infrastructure went to a halt. 3. Same place. The prod datacenter lost power, so we needed to failover to DR. "Let's ask the new guy to recreate 10k scheduling jobs in the DR env." Went surprisingly well, except that importing disabled jobs re-enabled them for some reason. An old env restart script kicked off on-schedule. 4. A new column caused my daily db import script to fail, which was only noticed after a few days of zero market data. 5. Overzealous find commands caused trades to fail (latency = bad pnl) 6. At an HFT firm, installing logstash included logstash-web which had a bad config that upstart continuously restarted. JVM restarts = bad news. 30k lost that day, apparently. 7. A typo in a script caused my cset shield script to bind the opposite cores. I fixed it the next morning after an angry wakeup call. Huge pnl improvement from this work, or it would've probably led to me being fired. I've seen: 1. A domain controller be brought to its knees after a typo'd password (bad authentication = no cacheing) from something similar to, "for i in hosts ; do sshpass $i hostname ; done. It took way too long to figure that one out. 2. Mid-day timezone change on every server. This was in clearing, so lots of backlash from clients here. 3. Plenty of accidental reboots. 4. DR failover scripts that have zero way of working. After complaining about this, management decided to task correcting that script to me (ugh). 5. 50 or so bad code releases. Devs, y'all aint in the clear ;) 6. Miraculously never saw anything bad from this, but an old company would require us to do backups on their prod databases by clicking through old school Solaris CDE dropdowns: right click on server -> backups -> create backup. We had to do this for about 30 database servers which were then used for testing over the weekend. The re-import was done the same way. 7. A windows admin ran an rsync with an incredibly shitty GUI on a production market data archive server with the "delete if non-existent" flag checked. We thankfully had a backup, but that backup would have taken 16 days to restore. I left after about day 8. 8. A server in a perf environment was brought over to prod by me. I recommended it be freshly wiped, as the number of unknowns (including user error) is so large that it's probably a time save to do so. Enough insistence of that forced me off that project, where my boss almost immediately and accidentally wiped the RAID. We were a week late in getting that server ready. (I couldn't help but grin) Point is, none of us were fired for any of the things we'd done wrong. It's hard to punish an accident, especially when the accidents stem from the folks before you, or bad management decisions. It truly is a mark of a good sysadmin when you've fucked up so badly that even non-techs say in disbelief, "Jesus.. whoops." edit: formatting
- sqldba 11y ago> I recommended it be freshly wiped If there's one thing I hate about the industry it's the adamant refusal in almost every single case to ever just "migrate a server" onto a fresh server. Every time there are known problems. Every time they get carried over. Every time there's 100,000 excuses not to do it. And in the end it's never worth it to avoid it. > It's hard to punish an accident Totally. I think identifying risk is super important. And if you identify it, and it's not cost/time effective to avoid it, and it gets approved - then you're off the hook for accidents.
- rpgmaker 11y agoWhen I was relatively new in my first job I forgot to include the WHERE clause in an update, essentially resetting the value for the entire table. Needless to say I felt awful and I was ready to hand in my resignation right after the issue was sorted out (I even printed my resignation letter). Luckily there was a relatively recent backup (not as recent as it should've been though... but I obviously wasn't the DBA) and things went back to normal relatively soon. Throughout the process the team shared their DB-related war stories with me. Everyone seemed to have had a similar experience happen at some point during their careers and knowing that made me feel a lot better. I ended up changing my mind and decided not to quit.
- hatter 11y agoMuch like the folk putting echos into find and wildcarded shell commands to check the output, I'll often start manual sql updates by doing 'SELECT something FROM database WHERE...' and check the output rows match my expectations before hitting the up arrow to replace the SELECT with an UPDATE. For bigger tables, I use COUNT(something) if I expect it to be long but I have an idea of rows affected, or LIMIT if that's going to give me an idea that it's doing the right thing.