4 ms·
Something not mentioned yet (at time of writing) is the importance of self-healing data. You can always just hard-restart an app every minute or so to be resil
by DelaneyM 4y ago
Something not mentioned yet (at time of writing) is the importance of self-healing data.
You can always just hard-restart an app every minute or so to be resilient to nearly any failure condition in runtime (not that I've _ever_ done that before...), but if the data gets into an invalid state you're stuck.
The one time I needed extreme resiliency and recoverability I used a write-only DB with a materialized view which updated on change or startup, and every write in a transaction. I also tailed the DB updates to a file on disk which replicated regularly off-site. It could automatically recover from nearly anything, and was remarkably easy to set up. The hard part was the materialized view, but I "needed" that anyways as I wanted to keep a full audit log as the primary db.
What constitutes resilient data is going to be unique to your use case of course, but consider resiliency from the DB up.
Also I suggest investing heavily in grokkable and relevant runtime observability. (Don't just emit inline comments, put some thought into relevant data and alerts.). Often you'll see a failure coming days ahead of it causing a problem, and you won't need the app to self-heal.
- pjc50 4y agoYes, this is exactly it. It's also an important part of the "crash only" concept - the moment that you detect a bug or inconsistency, it's important to crash so that you discard potentially corrupt state in memory and re-load from the last known good point. This makes it harder to write out wrong state, because that's much harder to recover from. The absolute nightmare scenario is state getting corrupted causing a crash-restart loop. (a couple of days ago we had the "everything is about state" discussion on here: this is why separating "persistent state" into a separate box and keeping a very careful eye on it is important)