3 ms·
How does this work? This means the consumers have to somehow know the IDs of every message that's been sent and if it's been processed. Either you have a share
by Kalium 4y ago
How does this work? This means the consumers have to somehow know the IDs of every message that's been sent and if it's been processed.
Either you have a shared database and all consumers are local, in which case why are you passing messages at all, or you have a distributed system somehow. If you have a distributed system, you've got this problem, and going recursive probably won't help much.
Or I've missed something.
- iainmerrick 4y agoIt depends what you mean by “processed”. If there’s some particular action you want to take in response to the message, then you put the database next to wherever that action takes place. If it’s a notification displaying on a phone, the phone holds a DB that tracks what notifications have been shown. You can’t reliably show a notification on exactly one of a user’s devices. But that’s no biggie; display it exactly once per device, remove it everywhere when acknowledged anywhere.
- Kalium 4y agoThat sounds like at-least-once delivery with recipient systems tracking known received messages. Which is to say faking exactly-once with the power of idempotency, rather than exactly-once. This is academic with enough abstraction, but when you're designing a system and implementing the work of processing a message it can be a pretty important distinction. Especially if you're past the point where the list of seen items becomes a chokepoint for synchronization between consumers. > You can’t reliably show a notification on exactly one of a user’s devices. But that’s no biggie; display it exactly once per device, remove it everywhere when acknowledged anywhere. Potentially messy. Now you have two distributed system messages - the initial notification and the ack.
- iainmerrick 4y agoThinking about it, we’re probably violently agreeing with just some differences in terminology. I’m arguing that exactly-once delivery is possible if the message receipt happens in a single place, but maybe others don’t see that as a distributed system at all. Potentially messy. Now you have two distributed system messages - the initial notification and the ack. Not a problem, both are idempotent! :)
- deleted 4y ago[deleted]
- Kalium 4y ago> I’m arguing that exactly-once delivery is possible if the message receipt happens in a single place, but maybe others don’t see that as a distributed system at all. I think that only holds if the sender is also in the same place. Otherwise there's the very real chance of a message getting lost, turning your system from exactly-once into at-most-once. At this point your sender, consumer, and messaging systems are one system, so it's probably reasonable to question if that's a distributed system.
- strken 4y agoI don't understand why "fake" exactly-once delivery is fake. If I make you generate a random UUID in your system, then store the UUID in a centralised place in my system and ignore duplicate messages with that UUID, and you keep retrying until you get an ACK, why is that not exactly-once delivery? Is it because the UUID is "only" random with some collision chance? (But then how come a sequential ID from your system wouldn't count?) Is it because my system needs to trust your system? (But why is trust a factor here?) Is it because my system needs a centralised database? (But our two systems are still distributed when you consider them together, right?) Is this a semantic argument over the meaning of "delivery", where I'm not allowed to impose requirements or check a database because the message has already been "delivered"? (But then why are we quibbling over semantics?) Is it because the message broker becomes stateful? (But why is that a constraint?) I think from looking at the article that this is about delivery within a finite number of retries, but that seems like the kind of problem where in the real world we just ring each other when a message has been retried 50 times over the course of two days.
- whakim 4y agoIt's fake because the deduplication isn't a property of the messaging system, it's a property of the system consuming the messages. "Your system" is providing a workaround (duplication) for at-least-once delivery, not the message system. I think it's also worth noting that your suggested workaround has actually recreated the problem it's intending to solve. If you mark the UUID as "received" as soon as you get a message, how do you deal with duplicates if the processing fails? If you mark the UUID as "received" when you're done processing, how do you deal with the possibility that you'll receive the message multiple times? This can get hairy very quickly.
- iainmerrick 4y agoThe solution to that last problem is a transaction log.
- strken 4y agoWe come back to the original argument, which is that at-least-once delivery of idempotent operations is how you represent things that should happen exactly once in a system, and in a saner world this could reasonably be called exactly-once delivery. Every distributed systems engineer knows what "exactly-once delivery" is asking for, and in plain English it's valid to conflate the two, but for some reason the field has decided to treat the phrase as an annoying semantic pit trap for the unwary. Want to add an ID to your transactions to make them idempotent? Well, even though your transactions are now recorded exactly once, that wasn't technically delivery! Gotcha!
- lisper 4y ago> This means the consumers have to somehow know the IDs of every message that's been sent No, they only need the ids of messages they have received.