3 ms·
> So nodes are just to assume that other nodes are alive/reachable/functioning? How would you propose detecting node health without some form of heartbeat? gen
by rdallman 11y ago
> So nodes are just to assume that other nodes are alive/reachable/functioning? How would you propose detecting node health without some form of heartbeat?
generally there is probably enough traffic between nodes to just run failure detection off of txn failures / normal comms (rpc,etc) without running a separate failure detection mechanism, but this assumes every node is already talking to every node, which is not safe, but let's explore solutions to that. detecting failures from normal communication will likely be even faster than a heartbeat, provided traffic frequency is higher than the heartbeat (use case contingent). to get around the issue where all nodes aren't already in communication, could broadcast any local failures each node might observe and you save the traffic of the heartbeat in the normal case, a gossip mechanism could work to reduce the N^2'ness of that communication -- it seems likely that multiple nodes will detect a given failure at around the same time, we don't want each of them broadcasting that, potentially leading to more failures from increased net traffic. since the nodes who were in communication with the failed one have detected it, we really just need to update the nodes who weren't so that they can find out about it in a closed interval (same guarantee as a heartbeat, potentially looser bounds).
maybe this is being too optimistic, but at least it's fun to think about :)