3 ms·
Synchronous systems have the same problems you describe in your first paragraph, but worse (they tolerate less failures), so I'm not sure what you're trying to
by DblPlusUngood 10y ago
Synchronous systems have the same problems you describe in your first paragraph, but worse (they tolerate less failures), so I'm not sure what you're trying to say.
I don't buy the rest of your post for two reasons. First, partitions are not only problematic for writes. Externally consistent systems also cannot generally serve reads during a partition either. Furthermore, it is not true that partitions are not a realistic worry - i know of at least one large system that encounters partitions on a weekly basis.
Second, you seem to be arguing that for some reason it is easier to solve consensus at the application level. This is simply not true, otherwise we wouldn't have consistent databases.
- peterwwillis 10y agoI'm saying it depends on what you're doing... They may fail immediately, or they will simply survive implicitly because the operation was atomic, or it was a read operation and it didn't matter anyway, etc. Consistency is [sometimes] a lie. What are you gonna do, read from the database directly every single time a web server wants to use a user's session cookie? You cache it, you have cache controls, you try to invalidate caches or update them if a new operation changes the session, but there's totally a possible race condition that will be resolved if someone tries a write on an invalid session. There's plenty of valid harmless operations on stale data due to network partition, it's not the end of the world. Of course some designs are more vulnerable to network partition than others (some designs only work on one network, some require multiple networks) so there are of course cases where you need something like Paxos. I'm saying it's less common than people want to believe. I find higher level consensus easier because you have more context of what's going on. Session-based operations, for example; if the same operation at the same time is done by two different sessions, you can compare the timing of the operation, when each session was last created/updated, or simply kick back the operation to the sessions and inform them of the conflict and ask them to resolve. In any of these cases the application caught the conflict before it had to do a network communication hokey-pokey, and it can be programmed to make these decisions automatically, too. Letting Paxos decide might result in an immediate fix, but the users might not ever be informed of the consensus decision and one might be confused as to what the actual resulting operation was.