4 ms·
It's admirable when vendors don't try to abuse theory to make their product look better. There was some blog post last week that supposedly proved the impossib
by dondraper36 2y ago
It's admirable when vendors don't try to abuse theory to make their product look better.
There was some blog post last week that supposedly proved the impossibility of exactly once delivery wrong, but that was quite expectedly another attempt to rename the standard terms.
- convolvatron 2y agoif you employ persistence to keep sequence numbers, and/or retransmissions and redundancy, you can drive the likelihood of distributed consensus over time arbitrarily close to probabilities where it doesn't matter anyways (like the earth suddenly being coincident with the sun). so what are we arguing about?
- sethammons 2y agothere is a group of people that say that there is no such thing as exactly-once-delivery and they are correct. There is also a group of people who cannot hear nor read nor process the word "delivery" in that term and instead insist that exactly-once-processing is easily achievable via [deduplication, idempotent actions, counters, ...]. They are correct as long as they acknowledge "processing" vs "delivery" -- and there are several who refuse to acknowledge that and don't believe the distinction matters and they are wrong.
- rwiggins 2y agoOut of curiosity, how would you describe TCP in these terms? Does the TCP stack's handling of sequence numbers constitute processing (on the client and server both I assume)? Which part(s) of a TCP connection could be described as delivery?
- wmf 2y agoFrom TCP's perspective it has delivered the data when read() returns successfully. This is also the point at which TCP frees the packet buffers. From the app's perspective, TCP is an at-most-once system because data can be lost in corner cases caused by failures. (Of course plenty of people on the Internet use different definitions to arrive at the opposite conclusion but they are all wrong and I am right.)
- YZF 2y agoI agree this is what you'd consider delivery. Also agree everyone else is wrong an I am right ;) The similar question in TCP is what happens when the sender writes to their socket and then loses the connection. At this point the sender doesn't know whether the receiver has the data or not. Both are possible. For the sender to recover it needs to re-connect and re-send the data. Thus the data is potentially delivered more than once and the receiver can use various strategies to deal with that. But sure, within the boundary of a single TCP connection data (by definition) is never delivered twice.
- wellpast 2y agois there a real difference between delivery and processing? In a streaming app framework (like Flink or Beam, or Kafka Streams), I can get an experience of exactly once "processing" by reverting to checkpoints on failure and re-processing. But what's the difference between doing this within the internal stores of a streaming app or coordinating a similar checkpointing/recovery/"exactly-once" mechanism with an external "delivery destination"?
- ithkuil 2y agoYes there is a difference and it's quite crucial to understanding how two groups of people can claim seemingly opposite things and both be correct. It's important to note that by processing here we don't just mean any "computing" but the act of "committing the side effects of the computation"
- wellpast 2y agoA typical delivery target is a data store or cache. Writing to such a delivery can be construed of as a “side effect”. But if you accept eventual consistency, achieving exactly once semantics is possible. Eg a streaming framework will let you tally an exactly once counter. You can flush the value of that tally to a data store. External observers will see an eventually consistent exactly-once-delivery result.
- ithkuil 2y agoUltimately the behavior that matters is the effective behaviour of the system. But a system is built by many components. You can build a reliable system out of unreliable components. You can build a transactional system out of non-transactional components. However you pay a price for that, because you have to bear the consequences of this substrate in other layers, often in your business logic too.
- sethammons 2y ago> You can flush the value of that tally to a data store What if you can't flush? You can't guarantee a flush: the sender or receiver can get powered off or otherwise get stuck for an indeterminate amount of time. That flush is the delivery. The data, is it stored in memory? Then it was lost when the process was killed.
- andrewflnr 2y agoNot that I disagree with getting terms exactly right, but if the "processing" that happens exactly once in your messaging system is firing a "delivery" event to a higher level of the stack (over a local channel), that's observationally indistinguishable from exactly-once delivery to that higher level, isn't it? What am I missing?
- wmf 2y agoNope, because that event could also be lost/duplicated.
- Izkata 2y agoIn the systems I'm used to, that "passing up" is a function call within the same codebase. There isn't a second delivery system for a message to be lost/duplicated, though sometimes they obscure it by making it look more like such a system.
- sethammons 2y agoPower gets cut. Now what? A process can fail between two function calls. When your power is restored, what happens with that mid-processing request? Lost or retried?
- andrewflnr 2y agoOver a local channel? Like same machine or even same process? I mean, yeah, you might have a bug that messes up the local "processing", but that's not a deep truth about distributed systems anymore, it's just a normal bug that you can and should fix.
- ViewTrick1002 2y agoThe problems always stem from the side effects. You can achieve exactly once like properties in a contained system. But if the processing contains for example sending an email then you need to apply said actions for all downstream actions, which is where exactly once as seen by the outside viewer gets hard.
- mike_hearn 2y agoEmail has the Message-Id header which is basically an idempotency ID.
- SpicyLemonZest 2y agoWhat you can't do, and what naive consumers of distributed systems often expect, is drive the probability arbitrarily low solely from the producer side. The consumer has to be involved, even if they consider themselves to be a single-node process and don't want to bother with transactionality or idempotency keys.
- mgsouth 2y agoVendor: "Exactly once is possible. We do it!" No, as a vendor you're either deliberately lying, or completely incompentent. Let's insert the words they don't want to. "We do exactly once good enough for your application!" How could you possibly know that? "We do exactly once good enough for all practical purposes!" So are my purposes practical? Am I a true scotsman? "We do exactly once good enough for applications we're good enough for!" Finally, a truthful claim. I'm not holding my breath. Edit: More directly address parent: It's not possible to even guarentee at-least-once in a finite amount of time, let alone exactly-once. For example, a high-frequency trading system wants to deliver exactly one copy of that 10-million-share order. And do it in under a microsecond. If you've got a 1.5 microsecond round trip, it's not possible to even get confirmation an order was received in that time limit, much less attempt another delivery. Your options are a) send once and hope. b) send lots of copies, with some kind of unique identifier so the receiver can do process-at-most-once, and hope you don't wind up sending to two different receiving servers.
- sidewndr46 2y agoThis didn't stop AWS from at one point declaring "none of the above" as their delivery guarantees for S3 events: https://www.hydrogen18.com/blog/aws-s3-event-notifications-probably-once.html https://www.hydrogen18.com/blog/aws-s3-event-notifications-p... From what I can tell my blog post eventually motivated them to change this behavior