4 ms·
Consider the following sequence* of events: 1. client A reads from partition X 2. client A is unsatisfied with the network map, and requests another read from
by fallingsquirrel 2y ago
Consider the following sequence* of events:
1. client A reads from partition X
2. client A is unsatisfied with the network map, and requests another read from partition Y
3. (meanwhile) client B writes a value to partition X
4. client A reads from partition Y, sees the value is the same as partition X's (stale) value, and accepts the value as consistent
This is the same kind of behavior you might get if your servers used e.g. a buggy version of Raft or something. You can't get around the proof by just relabeling some of your server nodes as client nodes.
* In the spirit of distributed systems, I use this term loosely :)
- Dylan16807 2y agoThe value is now stale, but it was correct at some point between the read starting and the read finishing, right? That can happen with any read. Even without any partitions. Can't it? In that case I don't see the problem.
- lmm 2y ago> The value is now stale, but it was correct at some point between the read starting and the read finishing, right? Not necessarily. Maybe both versions of it were from partial writes that were never committed, so your invariants are violated (if we're talking about e.g. a credit account A and debit account B scenario). > That can happen with any read. Even without any partitions. Can't it? Depends on your isolation level. If your system has serializable transactions then it's supposed to give you a history equivalent to one where all transactions were executed serially, for example.
- Dylan16807 2y ago> Not necessarily. Maybe both versions of it were from partial writes that were never committed, so your invariants are violated (if we're talking about e.g. a credit account A and debit account B scenario). I'm pretty sure the scenario above is looking at committed writes. If you're reading uncommitted writes, you're not really in the market for consistency to being with. (Or you could be handling consistency by waiting to see if the transaction succeeds or fails, making sure it would fail if the data you read got backed out. But in that situation nothing goes wrong here.) > Depends on your isolation level. If your system has serializable transactions then it's supposed to give you a history equivalent to one where all transactions were executed serially, for example. Even then, a new write can happen before your "read finished" packet arrives at the client, making the read stale. Your entire transaction is now doomed to fail, but you won't know until you try to start committing it. For pure read operations, I'm not convinced it's a proper stale read unless the value was stale before the read operation started.
- lmm 2y ago> I'm pretty sure the scenario above is looking at committed writes. > If you're reading uncommitted writes, you're not really in the market for consistency to being with. But what does "committed" mean when you're only reading from one partition in a partitioned scenario? You literally can't tell whether what you're reading is committed or not (or rather, you have to build your own protocol for when a write is considered committed). > For pure read operations, I'm not convinced it's a proper stale read unless the value was stale before the read operation started. I think you can get a read that is half from a prior stale operation and half from a subsequent uncommitted operation, or something on those lines.
- Dylan16807 2y ago> But what does "committed" mean when you're only reading from one partition in a partitioned scenario? I would say you can't make new commits in that situation? I don't know, I didn't make up the scenario, I think you need to figure out your own answer and/or get clarification from fallingsquirrel if you want to talk about that kind of problem. > I think you can get a read that is half from a prior stale operation and half from a subsequent uncommitted operation, or something on those lines. What's the full timeline for that? If you're specifically talking about the ABA problem, that's trivial to fix with a generation counter.
- lmm 2y ago> I think you need to figure out your own answer and/or get clarification from fallingsquirrel if you want to talk about that kind of problem. Right. The point is that fallingsquirrel hasn't solved the hard part of the problem. > If you're specifically talking about the ABA problem, that's trivial to fix with a generation counter. Maybe, but you need to specify that that's what you're doing, and it may come with undesirable consequences.
- Dylan16807 2y ago> The point is that fallingsquirrel hasn't solved the hard part of the problem. Sure. My point is that I don't see any issue with the read they described, not that I think there is a solution to partitioning here.