4 ms·
> syncing of changes to/from object storage So you are suggesting to have both the "serverless" db and the storage? If so, you now have 3 problems.
by Existenceblinks 4y ago
> syncing of changes to/from object storage
So you are suggesting to have both the "serverless" db and the storage? If so, you now have 3 problems.
- vidarh 4y agoI'm describing capabilities that already exist. Reliably streaming sqlite data to object storage is not a new thing and supported by multiple separate implementations at this point.
- Existenceblinks 4y agoThe problem of having multiple sources of data is Consistency, Availability, Partition tolerance. It also depends on IO characteristic of apps. Who are readers, who writers, how long stale data is acceptable. Basically all the distribute system problems. In web app, frontend + backend + db sit next to it is enough of problem (e.g. SPA vs MPA state problem)
- vidarh 4y agoDid not at any point suggest multiple sources of data. EDIT: While there are things in their repos that suggests they might be thinking about moving towards allowing multiple writers, what's currently there suggests a single current active instance of each database, with the WAL being sync'ed to object storage so that in the case of failure and/or when doing a cold-start, the database can be brought back from object storage.
- Existenceblinks 4y agoOk sorry, I think I misinterpret your architecture. I think you mean something like HA with Litestream.
- vidarh 4y agoYes, similar to that. So you'd put up a proxy to handle auth, and match an incoming request either to a running database or to a cold database. If it's for a running database, it'd replicate to S3 or similar. If it's for a cold database, you sync from S3 or similar and start the server side process. To your point, you absolutely need to be able to reliably grant a lease of some sort to whichever frontend pulls down the database and starts and endpoint, or you're absolutely right you'll have huge problems. Absolutely won't be suitable for every kind of workload, but if you've already committed to running your stuff in a serverless setup, having your database(s) handled that way might be appealing.
- Existenceblinks 4y ago> but if you've already committed to running your stuff in a serverless setup Sounds good but I'm curious what's criteria (need) to architect like this in the first place.
- vidarh 4y agoLet's say you want to run a huge number of databases for customers; too much to run on an individual server. Now you have to shard. You can either try to split them between multiple MySQL/Postgres etc. servers, that are now each individually major risk factors, or you can design your system so you can just hook up more servers at will as long as the largest individual customer database can run on a single instance. I've run large numbers of Postgres databases, and it's not hard to automate, but it's hard to optimise for a setup where the usage patterns of individual databases are hard to estimate. Is your customer using it for batch jobs, or for persistent streams of data? Who can you colocate with whom? When the cost of shutting a database down one place and migrating it elsewhere becomes very low, this kind of scenario can potentially become a lot easier. In terms of from the user perspective, I'd expect you really wouldn't care, other than in terms of cold start times and cost. Except perhaps for batch jobs etc., where being able to write apps that checks state, obtains a lease, downloads the most recent version does it's job and uploads the result to durable storage might well be convenient rather than having to e.g. keep a bunch of databases constantly running.