5 ms·
PACELC is an extension of CAP, so it suffers from the same problem of trying to apply a universal clock to the whole system. If you do that, you will suffer the
by andras_gerlits 3y ago
PACELC is an extension of CAP, so it suffers from the same problem of trying to apply a universal clock to the whole system. If you do that, you will suffer these limitations. With client-centric consistency, you can work around these problems and global order "unfolds" in a "just in time" manner.
The consistency-levels can go all the way to SNAPSHOT.
- sausagefeet 3y agoCAP says nothing about a "universal clock over the whole system". CAP is about the decision that has to be made in some unit of the system, it could be the whole system or it could be a bit, at the point of an operation. It's physics, there is no way around it. You can make different decisions on the semantics your system needs, but if you have two nodes that physically cannot communicate but need to be consistent for a client to move forward, the client cannot move forward. Full stop. Could you please show a failure mode that this system can handle that CAP says is not possible?
- andras_gerlits 3y agoSure. Here: https://medium.com/p/5e397cb12e63#04a5 https://medium.com/p/5e397cb12e63#04a5
- sausagefeet 3y agoI'm not sure if you gave the wrong link or not but this link doesn't describe any failure modes and how OmniLedger allegedly resolves them.
- andras_gerlits 3y agoThis section discusses some failure modes (the first one is about the failure of a specific node): https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and-fulfil-the-promise-of-microservices-5e397cb12e63#373c https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and... This section discusses latency-spike mitigation (which is how Brewer defines CAP colloquially): https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and-fulfil-the-promise-of-microservices-5e397cb12e63#7df1 https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and... This section dissects the problem when trying to apply CAP to non-linearizable systems like SQL: https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and-fulfil-the-promise-of-microservices-5e397cb12e63#04a5 https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and... Again, if you're not happy with the lack of scientific rigour in this technical article (ie: not science-paper), you can connect the dots in this one: https://www.researchgate.net/publication/359578461_Continuous_Integration_of_Data_Histories_into_Consistent_Namespaces https://www.researchgate.net/publication/359578461_Continuou...
- sausagefeet 3y agoI've read all this and I saw no description of failure modes and operationally how they are resolved. "If a node disappears just replace it with a new one". Ok, how?
- mrkeen 3y agoTo be fair, I think it's fine to ignore some of the implementation details about restarting a failed node. You can probably assume some kind of replicated log that all distributed systems use. And you can also give the benefit of the doubt when allowing some number (less than the quorum) of nodes to fail (and letting them restart and catch up, etc.) while the system still makes progress ("CA mode"). After all, that's the point of distributing a system in the first place - there's no one master which can die and bring down the system. But yeah, at this point I think OP is just going to keep implying that partitions don't happen or something...
- mrkeen 3y ago>> if you have two nodes that physically cannot communicate but need to be consistent for a client to move forward, the client cannot move forward. Full stop. > This section discusses some failure modes (the first one is about the failure of a specific node): https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and https://itnext.io/how-simple-can-scale-your-sql-beat-cap-and... > In this setup, a copy of the node is replaceable by taking a (potentially days old) backup copy of it and replaying the events that happened since the time the backup was established. The replaced node does NOT have access to the events that happened after the partition. * If the replaced node serves up stale data anyway, you built an AP system. * If the replaced node refuses to serve up stale data, you built a CP system. * If you're pretending it has access to the latest events, there's no partition and you built a CA system.
- andras_gerlits 3y agoNo. CAP requires linearizability for its definition. If your consistency-model moves with the network's ability to communicate, your strong consistency can progress even if you somehow manage to lose your redundant replicas. This is what the CAP-section is about: https://medium.com/p/5e397cb12e63#04a5 https://medium.com/p/5e397cb12e63#04a5 This is the summary: "In other words: any distributed solution that fits the SQL standard can rightly claim that it scales SQL databases, and Brewer’s model can certainly accommodate a framework for that. His model however, is not the only kind of distributed SQL database that can exist, therefore his assertion that all distributed consistent systems must pick where they position themselves on his famous triangle is wrong. The system we explain here for example, is an exception. Formally: because our consistency model stays within the bounds of what the SQL standard allows and includes network communication; and informally because we can fine-tune latency variability according to the use-case of the specific datastore within the system and can even be reduced to only be a theoretical concern."
- tsimionescu 3y agoUltimately the problem we care about is very simple, and there is no way to solve it, despite your claims. Say we have a database of customers, with two replicas - one in the USA, one in Europe. Say a customer in Europe wants to update their shipping address. We ship products every month to this customer from the USA to their current shipping address, on the 10th of that month. The customer is updating their address on the 9th at 10 AM PST. However, we are in the midst of a massive network partition that started on the 9th at 02 AM PST and is expected to last until the 11th. Do we perform the update in the European replica and tell the customer it succeeded (giving up consistency)? If we do, the US side of the business will ship the products to the old address, even though the update happened a full day earlier. Alternatively, do we tell the customer the update failed (giving up availability)? If we do, then the customer can't even let the European side of the business know of their new address. This is the CAP theorem in a nutshell, and it is obviously inescapable. It doesn't require appeals to immediacy that are anywhere close to the bounds of relativity. And while 3-day long network partitions are quite rare, partitions that last for many hours are not.
- andras_gerlits 3y agoAny modification to existing data must be "haggled for" somewhere, you're right about that. When you say "partition event", what is being partitioned here? A specific communication-link between two nodes. It's entirely possible (no, extremely likely) that your EU node would have access to a different US node, but not the one that's having the partition-event right now. That's exactly the point of this section in the essay: https://medium.com/p/5e397cb12e63#7df1 https://medium.com/p/5e397cb12e63#7df1 Networks will have latency-spikes, but if you can stream time-information the same way you can stream others, you can use redundancies to mitigate the disruption of any single channel.
- mrkeen 3y ago> It's entirely possible (no, extremely likely) that your EU node would have access to a different US node CAP theorem beaten by declaring P to be unlikely.