3 ms·
This post is interesting overall. But saying the org needs to care about scalability and then not understanding sharding is an interesting vibe.
by samlambert 3y ago
This post is interesting overall. But saying the org needs to care about scalability and then not understanding sharding is an interesting vibe.
- dalyons 3y agoI think the author understands sharding very well. And the operational complexities of it and thus why it should be a last resort.
- samlambert 3y ago1. it's not as complex as stated. 2. sharding scales read heavy workloads just as well as write heavy.
- dalyons 3y agook lets look at it. Start with the absolute best case for sharding - a perfectly segregated data model, say either by user or customer id. No shared data, no joins between tenents. Some relatively straightforward framework code to redirect queries between N shards. cool. Operationally now you need N database clusters (at least a primary and a replica per cluster). You need to manage and monitor all those, figure cluster local and shard promotion schemes, figure shard based backups and restores. Probably figure out shard rebalancing. None of it rocket science, but a lot to get right, test and maintain. Now do all that operational stuff in a way that works with on prem installs (gitlabs primary customer type). And then add in the fact that practically noones data model is that cleanly shardable, so add support for cross shard joins, and global tables. Its pretty easy to see how a couple of read replicas (that are near zero cost operationally if you're using a cloud db) are a VASTLY simpler solution.