5 ms·
The problem of having multiple sources of data is Consistency, Availability, Partition tolerance. It also depends on IO characteristic of apps. Who are readers,
by Existenceblinks 4y ago
The problem of having multiple sources of data is Consistency, Availability, Partition tolerance. It also depends on IO characteristic of apps. Who are readers, who writers, how long stale data is acceptable. Basically all the distribute system problems. In web app, frontend + backend + db sit next to it is enough of problem (e.g. SPA vs MPA state problem)
- vidarh 4y agoDid not at any point suggest multiple sources of data. EDIT: While there are things in their repos that suggests they might be thinking about moving towards allowing multiple writers, what's currently there suggests a single current active instance of each database, with the WAL being sync'ed to object storage so that in the case of failure and/or when doing a cold-start, the database can be brought back from object storage.
- Existenceblinks 4y agoOk sorry, I think I misinterpret your architecture. I think you mean something like HA with Litestream.
- vidarh 4y agoYes, similar to that. So you'd put up a proxy to handle auth, and match an incoming request either to a running database or to a cold database. If it's for a running database, it'd replicate to S3 or similar. If it's for a cold database, you sync from S3 or similar and start the server side process. To your point, you absolutely need to be able to reliably grant a lease of some sort to whichever frontend pulls down the database and starts and endpoint, or you're absolutely right you'll have huge problems. Absolutely won't be suitable for every kind of workload, but if you've already committed to running your stuff in a serverless setup, having your database(s) handled that way might be appealing.
- Existenceblinks 4y ago> but if you've already committed to running your stuff in a serverless setup Sounds good but I'm curious what's criteria (need) to architect like this in the first place.
- vidarh 4y agoLet's say you want to run a huge number of databases for customers; too much to run on an individual server. Now you have to shard. You can either try to split them between multiple MySQL/Postgres etc. servers, that are now each individually major risk factors, or you can design your system so you can just hook up more servers at will as long as the largest individual customer database can run on a single instance. I've run large numbers of Postgres databases, and it's not hard to automate, but it's hard to optimise for a setup where the usage patterns of individual databases are hard to estimate. Is your customer using it for batch jobs, or for persistent streams of data? Who can you colocate with whom? When the cost of shutting a database down one place and migrating it elsewhere becomes very low, this kind of scenario can potentially become a lot easier. In terms of from the user perspective, I'd expect you really wouldn't care, other than in terms of cold start times and cost. Except perhaps for batch jobs etc., where being able to write apps that checks state, obtains a lease, downloads the most recent version does it's job and uploads the result to durable storage might well be convenient rather than having to e.g. keep a bunch of databases constantly running.
- Existenceblinks 4y ago> a huge number of databases for customers Is this one.db per customer? If so, how do you deal with schema migration?
- vidarh 4y agoCarefully ;) One option is to build your app to check for migrations on startup. I've run systems that'd do schema migrations on a per user basis for data stored in objects per user, and it worked just fine across a userbase of a couple of million accounts, on the basis that the only clients that connect directly also owns the schema, so nothing ever connects without immediately checking if a migration needs to run. But consider that for setups like this the database might well be the customers own database, so their schema might not be something it's your job to touch at all.