3 ms·
The three node minimum is to avoid split brains. You can still access the data with fewer (for recovery etc.) but the cluster is not HA. Read throughput will
by morgo 9y ago
The three node minimum is to avoid split brains. You can still access the data with fewer (for recovery etc.) but the cluster is not HA.
Read throughput will go up by being able to distribute reads amongst the cluster. Write throughput shouldn't really get better; since all nodes have a copy of the data. Thus; the maximum node count is 9.
On the last question, there is data size, and there is working set size (what needs to be in memory). You can actually stretch working set size by sending certain queries (i.e. reports) to one of the nodes and keep the others for more transactional queries, but for storage on disk - yes, it is a multiple of however many nodes you have.
But also consider: Much of the pain I've had in large DBs is not being able to keep enough backups on fast media for quick restore (i.e. I'd like to have every day for the last 2+ weeks). From that pain point though, it's not a multiple - you just need to pick one node to backup.