4 ms·
Flynn is designed to be self-healing, if a server goes offline it will recover immediately and automatically and no confirmed writes will be dropped. We use a
by Titanous 10y ago
Flynn is designed to be self-healing, if a server goes offline it will recover immediately and automatically and no confirmed writes will be dropped.
We use a combination of Raft for service discovery and leader election, along with multiple instances of all core services so host failures within the quorum tolerance should not impact availability. The design of the data appliances is carefully considered, you can read an overview here: https://flynn.io/docs/databases https://flynn.io/docs/databases
If something catastrophic does go wrong, like the power going off in the datacenter, there is code that can resurrect the cluster when enough hosts come back online.
We did have some issues with recovery that have been fixed over the past few months, I'm almost certain the issue that you ran into has been fixed. Data is never destroyed by Flynn, so it's possible to recover the database from disk even if Flynn isn't running.
The system is also designed to tolerate partial failures gracefully, so if for example service discovery fails, the router will fall back to a cache and HTTP clients should not notice that anything is wrong.