3 ms·
I haven't seen or done anything of this scale before, but I did have a very sobering moment while working on a large online retailers stack as a systems enginee
by vhost- 10y ago
I haven't seen or done anything of this scale before, but I did have a very sobering moment while working on a large online retailers stack as a systems engineer.
We were rolling out a new stack in another data center across the country and before replication went live, I decided to connect and check things out. Our chef work hadn't completed for the database hosts, so I decided to install some OS updates by hand using pssh on all the MySQL hosts and saw a kernel update. So I thought, the DC isn't live yet, no replication is running, I'll just restart these servers. So I did using pssh again and then I caught a glimps at the domain in some output and my face went completely pale. I restarted the production databases... all of them. And they all had 256GB of ecc memory. It takes a very long time for each of those machines to POST.
I contacted the client and said the maintenance page was my fault and was fully expecting to be fired on the spot, but they just grilled me about being careful in the future, and then laughed it off.
I've been the most careful ever since then. It scared me straight. Always make sure you are in the right environment before you do anything that requires a write operation.
- olig15 10y agoExactly the right response. You're not going to make that same mistake again, but if you were fired, your replacement very well might.