5 ms·
Read more closely. Durability was NOT a requirement. Since we've decided to scope out access to historical logs from the problem we were trying to sol
by cwp 11y ago
Read more closely. Durability was NOT a requirement.
Since we've decided to scope out access to historical logs
from the problem we were trying to solve and focus only on real-time log
consolidation, that feature of Kafka became an unnecessary penalty without
providing any benefits.
- ZielenMiejska 11y agoAnd that puzzles me: how useful is a logging solution which can loose log messages by design? In an environment where "babies will die"? Maybe there is durable logging somewhere else? What am I missing here?
- vidarh 11y agoConsider the situations where it's likely to lose log messages: When the system is overloaded. Exactly the type of situations where you don't want to slow the system down further by spending resources on trying to let non-essential services survive. Presumably their tradeoff is that it's more important for the system to remain available than for every log message to be delivered. Then secondly you try to deliver log messages with as high reliability as possible. Often it is better to design for non-essentially components to fail early, or at least prevent their resource usage from escalating and dragging down other parts of the system (in this case, fixed buffers in 0MQ lets them isolate load in one part of the system by simply locking the rate the drain the buffers at below a suitable threshold that's normally fast enough).
- acconsta 11y agoThe problem is that overloaded 0MQ pub sockets drop exactly the data you don't want to drop — the newest data. Maybe they handle that in the application layer, but it's not clear.
- PieterH 11y agoIt's a problem in theory, not in practice. The notion of throwing away old data (or coalescing new with old) is relevant to slow queues. It turns out to be irrelevant with ZeroMQ. First, message rates with ZeroMQ are often hundreds of thousands per second. The architecture must be designed so that no buffers, anywhere, overflow. If they do, you have a problem, usually a slow subscriber. Throwing out older data doesn't cure the problem. What ZeroMQ does is punish the slow subscriber by dropping so that the publisher doesn't crash. It's not recovery for the subscriber, it's protection for the publisher (and thus for other subscribers). Second, trying to delete old messages is complex and sometimes impossible (if they're already in system buffers). The design of ZeroMQ's internal pipes has one writer and one reader, without locks. For the writer to mess with the reader would slow down everything and introduce risk of bugs. Dropping new incoming data is the only way anyone has ever found to keep things running at full speed. These design choices were often delicate and counter-intuitive, yet they have turned out to be mostly accurate.
- acconsta 11y ago>trying to delete old messages is complex and sometimes impossible I'm not familiar with ZeroMQ's data structures, so forgive my ignorance. At the high water mark, why can't the consumer throw away old messages instead of the producer throwing away new messages? There are no locks or bugs — that's what the consumer does anyway. I'm not saying that should be the default behavior, but perhaps an option. It's cleaner than silently dropping new messages, then sending an entire buffer of old messages if the subscriber recovers.
- vidarh 11y agoI love the thinking behind ZeroMQ... I've been a big fan of the work you guys have been doing all the way back to Libero (I used it to clean up a blazing fast multiplexed NNTP server with it back in '96 or '97; the first version did all the state transitions manually)
- jacques_chester 11y agoThis is why you should drop metrics first, then logs. Also, as you've alluded to, guaranteeing log availability provides an enticing cross-tenant denial of service vector. Vomit enough log messages into a shared fabric and you can begin to affect your neighbours. ZeroMQ, if I read right, diminishes some of this risk.
- jacques_chester 11y agoI read it differently, based on: At Auth0, high availability is the high order bit. We don't stress as much about throughput or performance as we do about high availability. We don't have a concept of a maintenance period. We need to be available 24/7, year round. Given the nature of some of our customers, if we are down, "babies will die". Though, as you point out, they simply dropped durability from their requirements.