4 ms·
I am confused by the query that had the problem. Specifically, I am confused by why the sharding is done by user id. Even the largest Slack instance probably h
by Robin_Message 4y ago
I am confused by the query that had the problem. Specifically, I am confused by why the sharding is done by user id.
Even the largest Slack instance probably has under 100,000 users and less than 1000 peak messages per second. That feels to me like it could be served by a single master DB. It feels like that would be a better way to shard.
Major downside is the major difference in shard sizes so some management/migration might be needed but it seems doable to me.
Certainly it feels naively that scaling should be easy due to the way slack instances are independent (unlike say Twitter).
- sophiebits 4y agoWorkspaces are not completely independent, such as for the Slack Connect feature. Some more details here: https://slack.engineering/scaling-datastores-at-slack-with-vitess/ https://slack.engineering/scaling-datastores-at-slack-with-v...
- Robin_Message 4y agoThanks! Indeed, the variability of size and usage of each instance was the big issue, and then doing features that crossed instances meant they'd be crossing shards whatever they did, so it made sense to fix the variability issue. (I'm also surprised/reminded how fast Slack grew and how quickly it became effectively ubiquitous — I think every company I've contracted for in the last five years has used Slack).
- iamcal 4y ago> Even the largest Slack instance probably has under 100,000 users and less than 1000 peak messages per second. This is not true, by an order of magnitude.