4 ms·
The more-fool-proof option for any distributed storage system is to not deploy multiple package upgrades across your entire fleet without any testing. This ent
by tene 9y ago
The more-fool-proof option for any distributed storage system is to not deploy multiple package upgrades across your entire fleet without any testing. This entire problem would have been avoided (for any distributed storage system) if they had first tested the upgrade on a single server before rolling it across their entire fleet. If you've got different code sitting around on disk ready for a surprise upgrade whenever something eventually someday gets restarted, you're going to be in for a bad time.
> It turned out that all three systems had been updated a few times without restarting the OSD. No OSD could start anymore. We kept the last two OSD running (this turned out to be a mistake). The file servers, running Gentoo, also had a profile update done by another administrator.
If you don't have the resources to perform safe upgrades, the fool-proof option is to pay a vendor to run your storage for you. I agree that if you want to run a reliable distributed storage system, you need a professional sysadmin who has enough time to maintain it safely. I further agree that if you don't have any professionals who are funded to dedicate at least some time to this, you'll probably have a lot less failure by just running a big dedicated node providing NFS, iSCSI, SMB, or whatever.
I can't think of any software that I'd call fool-proof under "We've upgraded multiple versions without any testing, we have no systems in a known-good state, and we don't have any way to actually revert back to a known-good state".
- mjevans 9y agoA /lot/ of times the issue is one of budget. It's //really// nice to actually have a budget so you can setup a testing environment and actually validate the things you're about to do to your production environment.