3 ms·
Consensus protocols, durability and transactional semantics are (should be) closely coupled. I recall TigerBeetle discussing somewhere how they could achieve be
by refset 3y ago
Consensus protocols, durability and transactional semantics are (should be) closely coupled. I recall TigerBeetle discussing somewhere how they could achieve better throughput and durability guarantees by combining replication/recovery with the consensus protocol, instead of layering it above. I.e. disaggregating the log can be expensive. There's a reference in [0] that might elaborate.
> TigerBeetle is “fault-aware” and recovers from local storage failures in the context of the global consensus protocol, providing more safety than replicated state machines such as ZooKeeper and LogCabin
[0] https://github.com/tigerbeetle/tigerbeetle/blob/main/docs/DESIGN.md https://github.com/tigerbeetle/tigerbeetle/blob/main/docs/DE...
- shikhar 3y agoI believe TigerBeetle are alluding to their integration of protocol-aware recovery [1], which is a worthy consideration for the log implementation. Yet another engineering concern which can be offloaded. If the disaggregated log integrates some mechanism to support leadership "above" it [2], it can be functionally identical to a converged log. Efficiency-wise yes there will be some extra network messages – but networks are very high throughput [3] and fast (sub-millisecond within a cloud region) these days! [1] https://www.usenix.org/conference/fast18/presentation/alagappan https://www.usenix.org/conference/fast18/presentation/alagap... [2] https://maheshba.bitbucket.io/blog/2023/05/06/Leadership.html https://maheshba.bitbucket.io/blog/2023/05/06/Leadership.htm..., also on HN yesterday [3] https://blog.enfabrica.net/the-next-step-in-high-performance-distributed-computing-systems-4f98f13064ac https://blog.enfabrica.net/the-next-step-in-high-performance...
- refset 3y agoThanks, yes protocol-aware recovery was the context. Pretty sure I first heard it described in Joran's QCon London 2023 talk here: https://youtu.be/_jfOk4L7CiY?t=1460 https://youtu.be/_jfOk4L7CiY?t=1460 > If you want your distributed database to maximise availability, how your local storage engine recovers from storage faults in the write-ahead log needs to be properly integrated with the global consensus protocol.