5 ms·
I'm not distributed systems expert, but I don't see any attempt at responding to: > 1. Distributed locks with an auto-release feature (the mutually exclusive l
by codys 11y ago
I'm not distributed systems expert, but I don't see any attempt at responding to:
> 1. Distributed locks with an auto-release feature (the mutually exclusive lock property is only valid for a fixed amount of time after the lock is acquired) require a way to avoid issues when clients use a lock after the expire time, violating the mutual exclusion while accessing a shared resource. Martin says that Redlock does not have such a mechanism.
The section "Distributed locks, auto release, and tokens" talks about why Martin's solution isn't right (I'm not well-read enough to tell if that argument is correct), but never actually argues that Redlock has a mechanism to avoid violation of mutual exclusion due to timeouts.
Perhaps I'm missing some background here: does such a mechanism exist?
- zzzcpan 11y agoIt doesn't exist. I think Martin called this out because of the fundamental impossibility to guaranty consensus asynchronously, i.e. with auto-release.
- squeaky-clean 11y ago(tl;dr, I don't think he does) I'm also not a distributed systems expert, but this is how I interpreted the arguments. AntiRez is responding to the scenario where a client acquires a lock, hangs, a new client acquires a lock, and the old client comes back and tries to write something based off of a now stale lock. As demonstrated in this image [0] So the random token ensures that even though many clients may "think" they have a lock, only up to a single client can ever actually have one. The write in the diagram which corrupts the database shouldn't happen, because you'll need to check that your lock token matches the service. [1] (excuse the terrible MSPaint modifications). Which covers the common cases of lock timeout. But the clients aren't the only part of the system that can violate the mutex. Martin seems to have already thought of this rebuttal: You cannot fix this problem by inserting a check on the lock expiry just before writing back to storage. Remember that GC can pause a running thread at any point, including the point that is maximally inconvenient for you (between the last check and the write operation). [ ... ] If you still don’t believe me about process pauses, then consider instead that the file-writing request may get delayed in the network before reaching the storage service. Packet networks such as Ethernet and IP may delay packets arbitrarily And I don't believe AntiRez answers this (or I'm not understanding when they do :P ). Even if a client makes a write-request at a perfectly valid point in time, when they own the lock, how do I know that by the time the write-request reaches storage, that the lock is still valid? The difference boils down to "redlock guarantees requests are sent during a valid lock", but not "redlock guarantees requests transact during a valid lock" Martin's fix [2] is similar to the token method suggested, except: A) the write-if-token-matches logic is handled directly by the storage layer. B) It uses an auto-incrementing token. He spends a few paragraphs making the point that redlock cannot generate good fencing tokens, because it has no good auto-increment consensus. But I don't see how a random token is any worse than an incrementing one if you're just going to straight up reject requests where the token doesn't match. Antirez addresses this, without addressing A. I would also really like someone to correct me if I'm missing something. The Redlock implementations are very bare-bones with tests or examples of how to actually use them. The Ruby library has a more "advanced" example, which says this at the point where they try writing to the file system: # Note: we assume we can do it in 500 milliseconds. If this # assumption is not correct, the program output will not be # correct. Which doesn't give me good faith in this... [0] http://martin.kleppmann.com/2016/02/unsafe-lock.png http://martin.kleppmann.com/2016/02/unsafe-lock.png [1] http://i.imgur.com/8rezAaX.png http://i.imgur.com/8rezAaX.png [2] http://martin.kleppmann.com/2016/02/fencing-tokens.png http://martin.kleppmann.com/2016/02/fencing-tokens.png [3] https://github.com/antirez/redlock-rb/blob/master/example2.rb https://github.com/antirez/redlock-rb/blob/master/example2.r...
- zzzcpan 11y ago> But I don't see how a random token is any worse than an incrementing one if you're just going to straight up reject requests where the token doesn't match. No, random tokens are not worse per se, they just have no place in the algorithm Martin described. It's a completely different system. In that system incremental tokens presume a mechanism that gives you a way to have some order of events. And one can rely on that order to decide which requests with which tokens to reject.
- antirez 11y agoI don't address what you do when the mutual exclusion is violated, because I show how a different token for each lock is good enough to avoid the races that Martin worries about, but at the same time I'm deeply skeptical in real-world use cases you often have this luxury. Very very often distributed locks are used when you need the lock as the sole way to avoid race conditions, so IMHO is a purely theoretical arguments for most people that are going to use Redlock. When you can use the token, use it, and even consider of not using the distributed lock at all if not to gain performances by avoiding race conditions all the times you can.