2 ms·
Subscribe-on-update databases are, almost by definition, problematic to use at scale as a generic storage solution. The fundamental problem is that they do not
by remon 6y ago
Subscribe-on-update databases are, almost by definition, problematic to use at scale as a generic storage solution. The fundamental problem is that they do not solve many real world problems efficiently enough to warrant the significantly higher running cost. Of course there are exceptions but it'll be hard to launch a MongoDB type of product that uses a subscription only model (see Firebase RTDB and its problems and lack of adoption in this space)
The reason developers gravitate towards subscription/rx based paradigms is because it results in very clean architecture and code. Unfortunately it comes at a cost which increases rather than decreases per client/user when volume increases. Some companies or projects can absorb that cost but not all, and typically less so when the project or its userbase grows.
A subscription based model will do work whenever data changes for each subscriber whereas more traditional pull based architectures only do work when a client specifically needs the data. This can be mitigated to some extent by being micromanaging subscriptions but that kills most of the value of the model.
There are also plenty of issues with this model if multiple clients are allowed to write to the same data which every single example project seems to try and do. There's a reason master-master updates, consensus algorithms and CRDTs all come at significant cost. It's usually hard. And when it's easy you probably don't need it subscribe-on-update in the first place.
- joshribakoff 6y agoI think you’re oversimplifying it. I was on teams that participated in architecture design for this stuff at Twitch. If we had 100,000s of clients polling at the same moment, that would absolutely knock over the server, the canonical “stupid easy” fix is to add jitter, which can be done on either client or server. In fact the pull based approach is doing work all the time to process all the polling. A push based approach only does work when needed (when data changes). I’m not saying it’s not new or hard to scale. I’m just saying objectively that push based is more efficient at least in terms of raw data sent down the wire