3 ms·
> Secondaries cover this case for Gitlab, it seems, but that comes with a set of caveats as well (namely async availability of data). Another caveat is that al
by mslot 9y ago
> Secondaries cover this case for Gitlab, it seems, but that comes with a set of caveats as well (namely async availability of data).
Another caveat is that all secondaries will end up having more or less the same stuff in their cache. As your data set grows bigger that becomes unsustainable, because you're going to read from disk more and more.
When you shard across N nodes you can keep N times as much data in the cache. Combined with N times higher I/O and compute capacity, that can actually give you way more than N times higher read throughput than a single node (for data sets that don't fit in memory), and you can get much higher write throughput as well.
- sb8244 9y agoGreat points, thanks!