4 ms·
I'm curious if there's a use case or edge scenario I'm not aware of that noout does not cover. I can't say I've got exhaustive testing for this, but when perfo
by tene 9y ago
I'm curious if there's a use case or edge scenario I'm not aware of that noout does not cover. I can't say I've got exhaustive testing for this, but when performing maintenance on a ceph cluster I ran at a previous employer, "ceph osd set noout" before maintenance did exactly what it is documented to do and prevented rebalancing while OSDs were stopped. Despite my poor memory, I'm pretty confident about that, because at least one upgrade I performed did end up requiring some data migration (I believe it was migrating from straw to straw2 bucket type), and we only made that change after all nodes were successfully upgraded.
http://docs.ceph.com/docs/master/rados/troubleshooting/troubleshooting-osd/#stopping-w-out-rebalancing http://docs.ceph.com/docs/master/rados/troubleshooting/troub...
My memory is generally quite poor, but I vaguely recall this feature being present for as long as I've been familiar with Ceph. I obviously have no way of knowing what happened in that meeting you were in, and maybe the other people who were proposing using Ceph were not very familiar with performing maintenance on it, or maybe there was other constraint or use case they had in mind, but I'm quite confident in saying that the normal case for Ceph maintenance involves only a very marginal amount of data movement (to bring the temporarily-down OSD back up-to-date with changes that occurred while it was down).
- kjetijor 9y agoI've used ceph since firefly - and noout existed back then as well. Assuming that you run with 3 failure domains, and only maintain one failure domain at a time. Noout mostly gets the job done. What it doesn't do for you is save you from an actual failure in a different failure domain during maintenance. EC pools & k+(m>=2) or replication > 3 - would cover this as well. We've had mostly great success with noout + maintain failure domain at a time, wait for recovery, proceed to next failure domain, repeat until done. To the point where we've been comfortable leaving a lot of the babysitting & work to machines.