6 ms·
> However, PAR would mean that a cluster can at least recover itself and then replace the disk in the background, without this being a showstopper event. This
by _benedict 5y ago
> However, PAR would mean that a cluster can at least recover itself and then replace the disk in the background, without this being a showstopper event.
This isn’t a showstopper event, it’s a process stopper event. I also outlined non-PAR approaches to recovering just fine without replacing the node.
> Theoretically, but in practice
I also have extensive practical confirmation that this approach works well.
> This aspect of recovery is often simply not tested.
I agree that test suites do not cover this well enough, and I commend you for working this into Tiger Beetle. This is something I hope to expand Cassandra's new deterministic simulation framework to incorporate in future as well, but Cassandra does benefit from a great deal of real world exposure to this kind of fault.
> PAR also shows cases where even a single local disk sector is enough to bring down a cluster that doesn't implement protocol-aware recovery for consensus storage.
I’m not sure I agree. The paper discounts reconfiguration because there could only be f+1 live processes and one could have corruption in a relevant sector, but under normal models this is simply f live processes. It seems that to accommodate this scenario we must duplicate the record identifier, and store both separately from the record itself.
It's not clear to me this scenario warrants the additional complexity, storage and bandwidth, as I’m not sure guarding against this is enough to reduce your replication factor. But it's worth considering, and this particular recovery enhancement is quite simple (we just need to duplicate the ballot in Paxos, so we can arbitrate between a split decision of other replicas that retain an intact record), so thanks for highlighting it more clearly for me.
> this kind of fault is addressed by Viewstamped Replication's 2012 revision
Could you point me to the place in the paper? I cannot see how VR solves this problem without introducing a risk of inconsistency, without necessitating an additional round-trip before responding to a client. Specifically, I think there are only three ways to address this particular problem, and I don’t see them discussed in VR:
1. The coordinator may record the responses of each replica before acknowledging to the client, so that the loss of any a write to any one disk may be recoverable.
2. The coordinator may require k>f+1 responses before answering a client, so that we may tolerate the loss of k-(f+1) disks losing a write
3. The coordinator may require an additional round-trip to record the consensus decision before answering a client
I don’t see any of these approaches discussed in VR. I see discussion of recovering the log from other replicas, but this cannot be done safely if the replica is not itself aware that it has lost data.
- _vvhw 5y ago> This isn’t a showstopper event, it’s a process stopper event. Again, the PAR paper describes the specific examples where a single disk sector failure on a single replica can be a showstopper event — for the whole cluster, even precluding reconfiguration since the log might no longer be functional. I know I've been hammering this point (single local disk sector failure = global cluster data loss or else global cluster unavailability), but it's really why PAR is essential for a distributed database, and why it won best paper at FAST '18. > Could you point me to the place in the paper? Sure, just grep the 2012 VSR Revisited paper throughout for "recovery protocol" to follow the discussion thread through the paper. It's one of the four main sub-protocols in the VSR paper. Here's the link to the paper itself: http://pmg.csail.mit.edu/papers/vr-revisited.pdf http://pmg.csail.mit.edu/papers/vr-revisited.pdf The protocol is section 4.3 but it's good to see the surrounding discussion in the paper where it's referenced. We also discussed this aspect of VSR in some extra detail at Aleksey Charapko's DistSys Reading Group recently: http://charap.co/reading-group-viewstamped-replication-revisited/ http://charap.co/reading-group-viewstamped-replication-revis...
- _benedict 5y agoI’d just like to say in preface that I appreciate this back and forth, as text can feel more combative than intended. I have learned in the process. > Again, the PAR paper describes the specific examples where a single disk sector failure on a single replica can be a showstopper event Again, I don’t believe it does, but I may have missed it. If you could specify the particular scenario you imagine that would be great. The paper explicitly only rejects recovery in the scenario that there are already f failed processes, so we are talking about f+1 faults. In essence, AFAICT the paper says that you can obtain additional failure tolerance by treating corruption as a non-fail-stop failure, and my contention is that the necessary replication factor for surviving other faults confers sufficient protection against corruption, and that as a result you most likely cannot reduce your replication factor with this improvement. > Sure, just grep the 2012 VSR Revisited paper Thanks. I read Section 4.3 before writing my prior reply, and I do not believe it addresses the problem.
- 5y ago