4 ms·
Hey, thanks a bunch for the input. In my proposed scenario, I would have expected there to be two databases - one at company A and another at AWS (company B).
by aetherspawn 6y ago
Hey, thanks a bunch for the input.
In my proposed scenario, I would have expected there to be two databases - one at company A and another at AWS (company B).
I plan to give company A a preinstalled server and ask them to just plug it in at their premises. But the server won’t be big enough after 6 months of data, so I want to charge them a small rate to push their data to my AWS when it gets too full.
The technology to sync the two databases and determine which data to keep on premises and which to push off to AWS is what I am trying to identify.
Hopefully that clarifies
- brudgers 6y agoThe only way to know what data can be offsite is to track actual usage and access patterns and tune the system to that. It's a database tuning problem not a theory problem.
- ciprian_craciun 6y agoUnfortunately there is no "out-of-the-box" solution for what you are trying to do. It depends on the actual database (or storage solution) that you actually employ; for example: * in case of Postgres there is streaming replication; but I think by default (and perhaps there isn't a way to disable that) it also "streams" `DELETE` operations; so you still have to do things manually; * in case of CouchDB (a document store) you could use replication, and then on the "on-premise server" (i.e. in company A) delete the old data and filter this operation in the replication; However, in the end I bet you'll end-up implementing a custom replication procedure. Which is not impossible, but not quite easy when you take into account network reliability. ---- A "quick and dirty" solution would be this: * dump data (either in SQL, JSON, TSV or whatever fits your use-case) in batches, of say 1 day or 1 week per batch; compress that with `zstd` or `lzip`; (depending on your data it should compress quite nicely to at least 75%;) * upload that to S3 for archival or staging area; * if you need on-line access to the data, import it in your database on AWS side;