5 ms·
Do you have a better solution? Genuinely interested since I'm about to implement webservice data change event notifications using probably RabbitMQ.
by nugator 11y ago
Do you have a better solution? Genuinely interested since I'm about to implement webservice data change event notifications using probably RabbitMQ.
- eropple 11y agoIt does require some up-front work, but I've generally moved over to Kafka for most of my stuff. Rabbit has some nice aspects, but as noted its HA options start at "dire" and escalate smoothly to "oh dear god no". Throughput with Kafka is very, very good and,in my experience, it's remarkably difficult to kill. If you want fully ephemeral topics, you can write the broker data to a tmpfs. But I really like keeping data around, because I can replay my topics later both for DR and for debugging.
- slowmovintarget 11y agoKafka makes reliable durable messaging not just possible, but solves all the usual attendant problems you try to design around. * Durable messages cause a slow-down when publishing, but not with Kafka because it uses the Linux kernel's page cache to write. * Durable messages are slow to read if they have to come from disk. Kafka is optimized to load sequential blocks into memory an push them out through a socket with very few copies. This makes for very fast reads. * Slow consumers can bog down the broker. Kafka stores all messages and keeps them on the topic for a time horizon. No back-pressure from slow consumers. * Disconnected but subscribed consumers cause messages to back up on disk and eventually clog the broker. Kafka stores all messages. There's no clogging or backup, that's just how it works. * Brokers must track whether a consumer actually received the message, failures can cause missed messages or clogs. Kafka clients may read from a given point on the topic forward. If they fail during a read, they just back up and read again. The messages will be there for hours/days/weeks as configured. With a rock-steady durable messaging system based on commit logs, all of those problems that arose from attempting to avoid durable messaging go away. Now you build microservices that emit and respond to events. Microservices that can "rehydrate" their state from private checkpoints and topic replays. And all of this with partition tolerance and simple mirroring. Although it isn't usually necessary, if you want, you can make all of that elastic with Mesos, too. Further reading: [1] https://engineering.linkedin.com/distributed-systems/log-what-every-software-engineer-should-know-about-real-time-datas-unifying https://engineering.linkedin.com/distributed-systems/log-wha... [2] http://mesos.apache.org/ http://mesos.apache.org/
- lobster_johnson 11y agoKafka is very good for synchronizing data streams, but I can't imagine that it's suitable for RPC?
- slowmovintarget 11y agoYou don't (or shouldn't) do ordinary RPC over messaging any way. Rather, use an Event Sourcing style: http://martinfowler.com/eaaDev/EventSourcing.html http://martinfowler.com/eaaDev/EventSourcing.html
- lobster_johnson 11y agoI didn't read that web page carefully, but it seems to describe a transaction log, which is what Kafka excels it, but it has precious title to do with RPC. RPC is point-to-point communication based on requests and replies. Kafka's strictly sequential requirement would be terrible for this because a single slow request would hold up its entire partition — no other consumer would be able to process the pending upstream events. Kafka is also persistent (does it have in-memory queues?), which is pointless for RPC.
- eropple 11y agoMessage queues, period, aren't particularly good for RPC. HTTP, as an online protocol, has such huge advantages that trying to replace it doesn't make sense to me. However, for comms between services where a consumer isn't waiting on the other end, a message queue is plenty appropriate--and Kafka is much, much better at that than RabbitMQ is in terms of throughput and data sanity. I also quite like NATS, but Kafka provides similar performance characteristics in the general case (generally higher latency being the exception, though I have never encountered latency-sensitive processes where a message queue made sense in the first place) and means babysitting fewer systems.
- lobster_johnson 11y agoTo be clear, NATS is not a heavy messaging broker like RabbitMQ. For one, it's in-memory only, and queues only exist when there are consumers: If you publish and there are no subscribers, the message doesn't go anywhere. NATS is closer to ZeroMQ than RabbitMQ or Kafka. A lot of people use HAProxy to route messages via HTTP to microservices — what's HAProxy if not a glorified message queue? If you don't use an intermediate — meaning you to point-to-point HTTP between one microservice and another — you have to find a way to discover peers, perform health checks, load-balance between them, and so on. Which you can do — services like etcd and Consul exist for this — but using an intermediary such as NATS or Linkerd [1] is also a great, possibly simpler solution. [1] https://linkerd.io/ https://linkerd.io/
- StreamBright 11y agoKafka saved my "life" few times. We had TTL set to 168h and somebody pushed a change to production that silently ignored a type of message. We realized it few days later. Luckily we could re-play all of the messages after fixing the code. I know there are so many things wrong with this, yet, Kafka is excellent at storing data for medium terms and that can be a real bliss.