4 ms·
> Additionally, if there is a delay or error causing a row number to be missing, we want the queue to pause processing until the missing message arrives before
by programd 3y ago
> Additionally, if there is a delay or error causing a row number to be missing, we want the queue to pause processing until the missing message arrives before sending the subsequent messages to the service
There's your problem. On the real Internet after some surprisingly finite amount of time the probability of some message never arriving is ~1.0.
The number one falsehood programmers believe about networking is that the network is reliable.
Even Google and Amazon, as good as they are, can't guarantee perfect network connectivity. In the real world you'll get hit with hardware failures, human failures (e.g. misconfiguration, bugs), random failures (flooding, cosmic rays), malicious failures (hackers), and so on. With this requirement you're guaranteed to have your system grind to a halt eventually.
Can you engineer around it? Well yes to a point, redundency and all that, but it's expensive and will still fail eventually. So both Google and Kafka gave you the correct answer.
- HealthLab 3y agoThanks for those thoughts. Our approach is that the sending system sends a numerical order for the messages. We took this approach in healthcare and it helped for when reliability of queues needed to be 100%. In that way we always knew if any message is missing in the queue. Given how often developers need to process a stream of messages - and assure the order of processing is reliable - it seems like such a surprise that queues would not have that native functionality for reliability.
- ac2u 3y ago>We took this approach in healthcare and it helped for when reliability of queues needed to be 100%. You might need to revisit what the concept of reliability means here. Not that it's possible, but your queuing and network system could be infallible but it'll not help if a producer of messages crashes and causes a missing message to begin with. >Given how often developers need to process a stream of messages - and assure the order of processing is reliable I mean what do you do if there's not a message there? Wait? How long? What if it never comes, does your queue just grow and grow in the meantime? Are you expecting to have these messages processed within the hour generally? What if it takes a month to get your missing message and then your system processes them in order a month later, is that ok? As a previous poster said, you can engineer around some things, but all you can do here is pull X messages from the queue/log, check if they have missing messages, if everything is present, process them in order, if it's not, wait and try again Y minutes later. But you still need to decide what constraints to relax, how many times you retry before you give up etc. There's no magic wand here that will provide a free lunch in the form of a product, the decisions have to be made. Also, a good way to reverse this problem is to ask yourself why you need the messages processed in order. If the reason is that certain data is computed when the messages arrive in and the data will only be correct if processed in order, then your answer might be to make the data derived computations from a table where you store the messages after consuming. That way you're not adding massive complexity and headache to your queueing infrastructure. (Kafka/Kinesis is another option here instead of a table, but most people don't need the throughput guarantees).