3 ms·
True, I imagined while writing the above that people will be thinking of different use cases (or in this case pathological cases). It is very important to prof
by lemmsjid 6y ago
True, I imagined while writing the above that people will be thinking of different use cases (or in this case pathological cases). It is very important to profile your system and rationalize all the things it is doing. I have multiple times, for example, seen production systems accidentally bottlenecked by having the wrong level of logging set, such that the primary cost of the system is trace logging, rather than anything having to do with customer value.
I would argue, however, that it has little to do with the decision of whether or not to distribute your system. I've spent a lot of time dealing, for example, with bottlenecked single instance RDBMS instances that are handling load they shouldn't be handling. (For example those accidental recursive queries that are often a side effect of nice ORM abstractions) I totally agree with you, and have seen it happen, that people who do not understand the performance characteristics of their system can reach for a distributed solution before they've understood what their performance issue was. But I've also seen plenty of situations where they reach for bigger hardware for the same reasons.
I think deciding whether or not to distribute means taking a disciplined approach to projecting the business needs of particular data entities you'll be dealing with. For example, if you have a users table, and it will ultimately store every human in the United States, then, well, that is quite do-able on today's single instance RDBMS systems, and you can project the theoretical growth over time. And if you need read load and HA, then you can go a replication route, or at least look at that first before doing something that reduces the quality of your transaction handling, like sharding. And then once the system is in place, taking a disciplined approach to profiling and quantifying the costs of the different aspects of the system and justifying their business value. For example having run large scale recommender systems, a typical decision might be to degrade the quality of an algorithm if it means saving a tremendous amount of money on processing.