3 ms·
I've leaned into it: all application state is on ephemeral storage and its constantly replicated "off-site" to a NAS running minio. To accomplish this, I have
by lknuth 2y ago
I've leaned into it: all application state is on ephemeral storage and its constantly replicated "off-site" to a NAS running minio.
To accomplish this, I have restricted myself to SQLite as storage and use Litestream for replication. On start, Litestream reconstructs the last known state before the application starts. [Source](https://github.com/LukasKnuth/homeserver/blob/912cbc0111e44d044e5c7e2bbac5a8473df089cd/deploy/modules/stateful_web_app/main.tf#L73 https://github.com/LukasKnuth/homeserver/blob/912cbc0111e44d...)
It works very well for my workloads (user interaction driven web apps) but there are theoretical situations in which data loss can occurr.
- amluto 2y agoYou need to store data somewhere. If you’re willing to have that location be off-site and to pay for remote storage and possibly egress, and you’re okay with your local installation not coming up if the off-site storage is inaccessible and with the latency involved in bringing your on-site data back, fine. Of course, if you really lean in to a complete lack of local persistent state and you configure your network or some other critical service like this, good luck recovering from an upgrade and complete loss of ephemeral state.
- lknuth 2y agoI'm not sure I understand your second point. Yes, if the external replication target is unavailable, I can't bring my local service back online. Same goes for if the replication target becomes unavailable and I don't realize it, there is potential for a lot of data to be lost if the application restarts. For my personal usecase this is fine. I also have monitoring setup to look for just this case. It's a tradeoff between resilience and simplicity that works for some use cases - mine included.
- amluto 2y agoMy second point is that, in a setup of any complexity, a black start is nontrivial. You have a network, with routes, DNS config, maybe VLANs, maybe a whole SDN. If that’s down because the machines running it are trying to pull their own configuration over the network, it won’t come back up. You can get pretty far into the weeds with situations like this. Facebook supposedly got locked out of their own datacenter due to a network outage preventing the access control system from accessing whatever service it needed to allow anyone to open the door.
- lknuth 2y agoI see what you mean. For my case, the network is much simpler. I'm also fine if an unavailable replication target means I can't start an application. The upside of my solution is that there is no scheduling requirement on which node the PVC was initially created. There is also a certain guarantee that I have a working, recent backup of the application data. Starting from scratch every time is also basically a backup recovery operation. It gives me confidence that there is a recent backup which is restorable.