12 ms·
Kafka at the low end: how bad can it get?
- enether 2y agothankfully early access for KIP-932 is coming in 1-3 weeks as the 4.0.0 release gets published
- film42 2y agoFirst time I've heard of KIP-932 and it looks very good. The two biggest issues IMO are finding a good Kafka client in the language you need (even for ruby this is a challenge) and easy at-least-once workers. You can over partition and make at-least-once workers happen (if you have a good Kafka client), or you use an http gateway and give up safe at-least-once. Hopefully this will make it easier to build an at-least-once style gateway that's easier to work with across a variety of languages. I know many have tried in the past but not dropping messages is hard to do right.
- akshayshah 2y agoCouldn’t agree more - the most exciting thing about KIP-932 is how much easier it’ll become to build a good HTTP push gateway. Uber wrote a Kafka push gateway years ago, when it was considerably harder to do well: https://www.uber.com/blog/kafka-async-queuing-with-consumer-proxy/ https://www.uber.com/blog/kafka-async-queuing-with-consumer-...
- mensfeld 2y agoDo you mind explaining what you mean by not being able to find a "good Kafka client" for Ruby? There are pretty good bindings to librdkafka and frameworks like Karafka (https://github.com/karafka/karafka/ https://github.com/karafka/karafka/) that provide many functionalities, including a Web UI.
- film42 2y agoLate reply... I don't have my notes anymore on kafka client evals in ruby. When evaluating it was for a former employer. Karafka is very impressive though. Well done on the library.
- PhilippGille 2y agoTFA mentions it in the third paragraph: > Note: when Queues for Kafka (KIP-932) becomes a thing, a lot of these concerns go away. I look forward to it!
- enether 2y agoyep, I was mostly clarifying the timeline
- jszymborski 2y agoWhat do people recommend? Especially for low levels of load, that doesn't require that the dispatcher and consumer are written in the same language.
- giovannibonetti 2y agoRedis, SQLite or even a traditional DB like Postgres or MySQL can all do a better job than that.
- alexwebr 2y ago(Author here) RabbitMQ or AWS SQS are probably good choices.
- kgeist 2y agoWe use RabbitMQ, and workers simply pull whatever is next in the queue after they finish processing their previous jobs. I’ve never witnessed jobs piling up for a single consumer.
- sea-gold 2y agoNATS https://docs.nats.io/nats-concepts/overview/compare-nats https://docs.nats.io/nats-concepts/overview/compare-nats
- MuffinFlavored 2y agoNATS/WebSockets are good for 1 publisher -> many consumer (pubsub) RabbitMQ is good for 1 producer -> 1 consumer with ack/nack Right?
- esafak 2y agoNATS does many-to-many.
- Joel_Mckay 2y agoActually, I used RabbitMQ static routes to feed per-cpu-core single thread bound consumers that restart their process every k transactions, or watchdog process timeout after w seconds. This prevents cross contamination of memory spaces, and slow fragmentation when the parsers get hammered hard. RabbitMQ/Erlang on OTP is probably one of the most solid solutions I've deployed over the years (low service cycle demands.) Highly recommended with the AMQP SSL credential certs, and GUID approach to application layer load-balancing. Cut our operational costs around 37 times lower than traditional load-balancer approaches. =3
- xyst 2y ago> Each of these Web workers puts those 4 records onto 4 of the topic’s partitions in a round-robin fashion. And, because they do not coordinate this, they might choose the same 4 partitions, which happen to all land on a single consumer Then choose a different partitioning strategy. Often key based partitioning can solve this issue. Worst case scenario, you use a custom partitioning strategy. Additionally , why can’t you match the number of consumers in consumer group to number of partitions? The KIP mentioned seems interesting though. Kafka folks trying to make a play towards replacing all of the distributed messaging systems out there. But does seem a bit complex on the consumer side, and probably a few foot guns here for newbies to Kafka. [1] [1] https://cwiki.apache.org/confluence/plugins/servlet/mobile?contentId=255070434#KIP932:QueuesforKafka-Handlingbadrecords https://cwiki.apache.org/confluence/plugins/servlet/mobile?c...
- enether 2y agoEven if you used a 100% random strategy - what OP described can still happen. Matching the number of consumers can still produce an uneven result too, and OP clarifies that even if the worst case he laid out doesn't happen - in practice there still are idle workers. For the same reason, he likely doesn't want to have 16 workers at all times
- NovemberWhiskey 2y agoKafka for small message volumes is one of those distinct resume-padding architectural vibes.
- tstrimple 2y agoOr, as mentioned in the article, you've already got Kafka in place handling a lot of other things but need a small queue as well and were hoping to avoid adding a new technology stack into the mix.
- kvakerok 2y agoYou haven't seen the worst of it. We had to implement a whole kafka module for a SCADA system because Target already had unrelated kafka infrastructure. Instead of REST API or anything else sane (which was available), ultra low volume messaging is now done by JSON objects wrapped in kafka. Peak incompetence.
- kevinherron 2y ago> for a SCADA system for Ignition?
- Joel_Mckay 2y agoWe did something similar using RabbitMQ with bson over AMQP, and static message routing. Anecdotally, the design has been very reliable for over 6 years with very little maintenance on that part of the system, handles high-latency connection outage reconciliation, and new instances are cycled into service all the time. Mostly people that ruminate on naive choices like REST/HTTP2/MQTT will have zero clue how the problems of multiple distributed telemetry sources scale. These kids are generally at another firm by the time their designs hit the service capacity of a few hundred concurrent streams per node, and their fragile reverse-proxy load-balancer CISCO rhetoric starts to catch fire. Note, I've seen AMQP nodes hit well over 14000 concurrent users per IP without issue, as RabbitMQ/OTP acts like a traffic shock-absorber at the cost of latency. Some engineers get pissy when they can't hammer these systems back into the monad laden state-machines they were trained on, but those people tend to get fired eventually. Note SCADA systems were mostly designed by engineers, and are about as robust as a vehicular bridge built by a JavaScript programmer. Anecdotally, I think of Java as being a deprecated student language (one reason to avoid Kafka in new stacks), but it is still a solid choice in many use-cases. Sounds like you might be too smart to work with any team. =3
- rockwotj 2y agoThe kafka protocol is a distributed write ahead log. If you want a job queue you need to build something on top of that, it’s a pretty low level primative.
- atmosx 2y agoWhy does everybody keep missing this point? I don’t know.
- nobleach 2y agoThere's a wonderful Kafka Children's book that I always suggest every team I work with read: https://www.gentlydownthe.stream/ https://www.gentlydownthe.stream/ The way I describe Kafka is, "an event has transpired... sometimes you care, and choose to take an action based on that event" The way I describe RabbitMQ is, "there's a new ticket in the lineup... it needs to be grabbed for action or left in the lineup... or discarded" Definitely not perfect analogies. But they get the point across that Kafka is designed to be reactive and message queues/job queues are meant to be more imperative.
- stickfigure 2y agoYour two-sentence description is excellent. That book, not so much.
- nobleach 2y agoI suppose that's fair.
- mumrah 2y agoNot for long. An early access version of KIP-932 Queues for Kafka will be released in 4.0 in a few weeks. https://cwiki.apache.org/confluence/display/KAFKA/KIP-932%3A+Queues+for+Kafka https://cwiki.apache.org/confluence/display/KAFKA/KIP-932%3A...
- brunoborges 2y agoFor a small load queueing system, I had great success with Apache ActiveMQ back in the days. I designed and implemented a system with the goal of triggering SMS for paid content. This was in 2012. Ultimately, the system was fast enough that the telco company emailed us and asked to slow down our requests because their API was not keeping up. In short: we had two Apache Camel based apps: one to look at the database for paid content schedule, and queue up the messages (phone number and content). Then, another for triggering the telco company API.
- araes 2y agoHaving never actually used this platform before, does anybody know why they named it Kafka, with all the horrible meanings? Per Wiktionary, Kafkaesque: [1] 1. "Marked by a senseless, disorienting, often menacing complexity." 2. "Marked by surreal distortion and often a sense of looming danger." 3. "In the manner of something written by Franz Kafka." (like the software language was written by Franz Kafka) Example: Metamorphosis Intro: "One morning, when Gregor Samsa woke from troubled dreams, he found himself transformed in his bed into a horrible vermin. He lay on his armour-like back, and if he lifted his head a little he could see his brown belly, slightly domed and divided by arches into stiff sections. The bedding was hardly able to cover it and seemed ready to slide off any moment. His many legs, pitifully thin compared with the size of the rest of him, waved about helplessly as he looked." [2] [1] Wiktionary, Kafkaesque: https://en.wiktionary.org/wiki/Kafkaesque https://en.wiktionary.org/wiki/Kafkaesque [2] Gutenberg, Metamorphosis: https://www.gutenberg.org/cache/epub/5200/pg5200.txt https://www.gutenberg.org/cache/epub/5200/pg5200.txt
- InDubioProRubio 2y agoBecause its a process: https://en.wikipedia.org/wiki/The_Trial https://en.wikipedia.org/wiki/The_Trial
- op00to 2y agoJay Kreps liked Kafka’s writing.
- denkmoon 2y agoNominative determinism.
- snotrockets 2y agoIt was named so based on the Idea is that like the author (who the term "Kafkesque" is coined after), Apache Kafka is a prolific writer.
- kod 2y agoKafka wrote a lot, and destroyed most of what he wrote. Seems like a good name for a high-volume distributed log that deletes based on retention, not after consumption.
- voodooEntity 2y agoWe build an Infrastructure with about 6 microservices and Kafka as main message queue (job queue). The problem the author describes is 100% true and if you are scaled with enaugh workers this can turn out really bad. While not beeing the only issue we faced (others are more environment/project-language specific) we got to a point where we decided to switch from kafka to rabbitmq.
- techcode 2y agoWhat that post describes (all work going to one/few workers) in practice doesn't really happen if you properly randomize (e.g. just use random UUID) ID of the item/task when inserting it into Kafka. With that (and sharding based on that ID/value) - all your consumers/workers will get equal amount of messages/tasks. Both post and seemingly general theme of comments here is trashing choice of Kafka for low volume. Interestingly both are ignoring other valid reasons/requirements making Kafka perfectly good choice despite low volume - e.g.: - multiple different consumers/workers consuming same messages at their own pace - needing to rewind/replay messages - guarantee that all messages related to specific user (think bank transactions in book example of CQRS) will be handled by one pod/consumer, and in consistent order - needing to chain async processing And I'm probably forgetting bunch of other use cases. And yes, even with good sharding - if you have some tasks/work being small/quick while others being big/long can still lead to non-optimal situations where small/quick is waiting for bigger one to be done. However - if you have other valid reasons to use Kafka, and it's just this mix of small and big tasks that's making you hesitant... IMHO it's still worth trying Kafka. Between using bigger buckets (so instead of 1 fetch more items/messages and handle work async/threads/etc), and Kafka automatically redistributing shards/partitions if some workers are slow ... You might be surprised it just works. And sure - you might need to create more than one topic (e.g. light, medium, heavy) so your light work doesn't need to wait for heavier one. Finally - I still didn't see anyone mention actual real deal breakers for Kafka. From the top of my head I recall a big one is no guarantee of item/message being processed only once - even without you manually rewinding/reprocessing it. It's possible/common to have situations where worker picks up a message from Kafka, processes (wrote/materialized/updated) it and when it's about to commit the kafka offset (effectively mark it as really done) it realizes Kafka already re-partitioned shards and now another pod owns particular partition. So if you can't model items/messages or the rest of system in a way that can handle such things ... Say with versioning you might be able to just ignore/skip work if you know underlying materialized data/storage already incorporates it, or maybe whole thing is fine with INSERT ON DUPLICATE KEY UPDATE) - then Kafka is probably not the right solution.
- techcode 2y agoThe other thing that's PITA with Kafka is fail/retry. If you want to continue processing other/newer items/messages (and usually you do), you need to commit Kafka topic offset - leaving you to figure out what to do with failed item/message. One simple thing is just re-inserting it again into the same topic (at the end). If it was temps transient error that could be enough Instead of same topic, you can also insert it into another failedX Kafka topic (and have topic processed by cron like scheduled task). And if you need things like progressive backing off before attempting reprocessing - you liekly want to push failed items into something else. While it could be another tasks system/setup where you can specify how many reprocessing attempts to make, how much time to wait before next attempt ...etc. Often it's enough to have a simple DB/table.