8 ms·
Kafka as an Antipattern
- sdfghswe 3y agoDeceptively deep topic, and yet such a superficial post.
- wolfi1 3y agook, pretty embarrassing I thought of Kafka the author first and wondered why he would have been an antipattern ...
- no_wizard 3y agoSeems like a lot of what I read about Kafka really makes it sound like using it is quite, well, Kafkaesque Why do so many engineers end up having such a struggle with an event sourcing system, yet the system itself remains highly popular I don’t know. I theorize the following: - Its flexible enough to do things like receive events (messages) and sending downstream events from those received - it can ingest events fast. A well tuned instance is very fast and can handle a lot of volume - it often allows a middleware log point (or other work types) for things happening throughout your whole system Perhaps all of these things (and more) are hard to attain using a different technology
- marcinzm 3y agoThe simple approach mentioned in the article gets annoying if you have micro-services that don't share a DB. You could add a shared DB or a nosql DB but then you may as well just add Kafka. Of course the key question then shouldn't be kafka or not-kafka but if you over engineered on micro-services.
- quickthrower2 3y ago- developer tries to get those jobs paying $50k/y more “what kafka, event sourcing and microservices experience do you have”
- mrkeen 3y ago> Why do so many engineers end up having such a struggle with an event sourcing system, yet the system itself remains highly popular I don’t know. It's the mental model that's simple: some services write down what happened, other services 'do their own thing' with that information. I can write the core of the system in 2022, with events like 'Joe wants to buy a bike', 'Joe owes us $200', 'Joe paid us $200', 'Send Joe the bike', etc. In 2023 I want to build the book-keeping service, in 2024 I want to build the inventory management system, and in 2025 I want to hook it up to a CRM and see if we can try to sell some bike parts to Joe. Why couldn't I just use REST for that? Because the recipients of the rest calls didn't exist yet. The bad part of Kafka (in my opinion) is how opinionated the consumer logic is (oftentimes by necessity, because of the whole distributed system thing). Sometimes I just wanna ask "what offset are you up to?", but end up in API hell, and am unable to do it.
- convolvatron 3y agoI agree that Kafka is alot of machinery for a fairly limited gain. but I really dislike this whole notion of 'antipattern', as if we can look at thing and assign a decontextualized thumbs up or down. that building systems is just a matter of assembling the right patterns, and avoiding the antipatterns.
- BoorishBears 3y agoSometimes a technology or pattern can be poisonous in a very specific way that warrants a label: when they're most alluring to those least equipped to leverage them. Just like microservices, you cross the activation energy to want Kafka very easily because it's appealing on résumés, sounds like a hedge against scale,etc. But there's a huge asymmetry in understanding the drawbacks to them. When you spin up these systems, the drawbacks don't hit you immediately, it feels like they're solving the problem you had, and it's not until you've invested immense amounts of sweat capital (and literal capital) that you discover how badly you screwed up. You need some way to match that low effort value prop with a low friction warning: this is not a panacea for your problems. It only seems simple, it's not simple, it will hurt you unless know it will hurt you and simply have the resources and scale to play through that hurt. — To me that warning is what's implied by "antipattern", it's not never use this, it's never use this unless you know why you should never use this.
- jshen 3y agoI agree completely. Kafka, and event based architectures, are highly overused, but are also the right thing sometimes. It's FAR easier to manage an architecture where systems call into the source of truth for a given piece of data via APIs, and you should only switch to an event model if you truly need to.
- deleted 3y ago[deleted]
- davewritescode 3y agoThe anti-pattern here isn't Kafka, it's using Kafka for 7,500 messages/day. Making your whole system asynchronous for that level of load is the textbook definition of over-engineering.
- pjscott 3y agoYes. For perspective, that's about one message every ten seconds.
- alexchamberlain 3y agoSorry for the naivety/not obvious from the comments: is that too much or too little? (I've used RabbitMQ much more than Kafka.)
- yCombLinks 3y agoKafka is designed to maximize scalability, millions of messages a second. It's a pain in the neck to manage if you don't need it
- esafak 3y agoToo little to need Kafka.
- throw1234651234 3y agoWay too little to justify event-driven architecture (let alone Kafka specifically), unless you have some specialized need like very slow event processing and need to display a "message received" notification to user before the processing happens. Or you really need the retry functionality and can't handle it some other ways. Most businesses have no (hue hue) business doing event driven architecture. There is way too much overhead for local testing and overall complexity, especially when you want to properly handle errors. "But, every developer should be able to set up their local." Yea, great, explain to the manual QA who may be amazing but just started 3 months ago.
- 3y ago
- chrsig 3y agoIt seems a lot of the complaints weren't about kafka itself, but rather seemed to stem from internal communication problems. Custom kafka message headers could very well be custom http headers, and the problem is the same. Kafka is just coincidental. Looking at the volume though, kafka is overkill. They most likely could have just used the database and reaped the benefits of doing everything in a single transaction, with easier row level locking. The post acknowledges this. I do think it highlights the need for a small scale kafka, though. It's conceptually great to have everything work off of logs, but kafka does add a non trivial operational burden.
- imachine1980_ 3y ago>small scale kafka, though. It's conceptually great to have everything work off of logs, but kafka does add a non trivial operational burden. does something like that exist ???
- chrsig 3y agoNot that I'm aware of. I've been very tempted to write my own.
- ssrc 3y agoDepending on the meaning of "small-scale kafka", both RabbitMQ and redis do support streams.
- chrsig 3y agoOne of my desires would be for it to be persistent. Hopefully with the option of different storage tiers, so as logs became older they could be moved to less costly medium and transparently fetched when requested. Having an event sourced system doesn't make much sense unless you maintain messages from the start of the system. You can snapshot state and resume in order to quickly rebuild from a known good state. That doesn't help if there was a logic error corrupting every state from the start, and a full rebuild is required. I'm unsure how redis streams behave with regard to cache eviction, nor am I familiar enough with rabbitmq to comment on it's behavior. It's been 10 years since I used either, and at the time neither were good solutions for a log based system.
- tetha 3y agoThis is what I've been starting to think about on a more abstract level: Introducing a new technology, a new system into a design isn't like putting a piece into a jigsaw puzzle, or even worse, trying to mold and force the system to fit whatever hole your design has. Many more specialized systems - and Kafka is one of them - should solve some problem, but they should also change your mental model of the system and you should look for the easiest way to introduce these heavy hitters. For example, if you use Kafka or streaming solutions like Flink or Spark, you should change your mental model to (possibly large), (possibly resplayable) streams of events and look for simple ways to get these event streams going and good ways to consume them. And then you need to let the design push you where it wants you to go. Like, at work, we recently had a discussion how it was so storage-expensive for a project to store all events of a day and how the query to count all of these events per tenant was taking so long. While they are using a streaming event processor in front of it. Like, what the hell - think in streams, tally up these events on the fly and persist that every hour?
- bosky101 3y agoImmediately starts to doubt OP's assumption/implementation of the "where it works great" Jokes aside, agree with others. For the 7500/day, I would just push these into an S3/minio folder. And then dequeue 100 or N once every 10/30/60/T seconds. play around for the right N,T. Then again am sure there maybe reasons/context/constraints unaware to us. Eg - the ingestion is spiky, with the possibility of all 7500 in a few secs/minute, you would want to first make sure that the http traffic can scale before getting to the point where it can actually connect and push to the queue Possible reason#2 - an intern who just finished up their first Kafka task; just got freed up #3 - or this was the only infra available and a choice had to be made with the time available at hand #4 - or this was an experiment to see for yourself A majority may not agree with your view, but none of us really are in your shoes. So I applaud you for sharing your thoughts anyway.
- aeturnum 3y agoThe way I think about this kind of problem is to remember that tools built to deal with huge scaling problems are generally dealing with a very complex set of variables. The tool is going to be designed to let you choose between all of those variables. There's no magic - just configuration whose complexity better matches that of your problem. That being said, if you are not yet in a situation as complex as the one your tool is designed to deal with, there is a very good chance you will waste some time starting to use such a tool "early." You might get that time back later when you scale, you might have the right people to set up the complex tool the right way for your simple situation, but you are taking a bit of a risk. As long as you go into the situation with your eyes open I think most people end up ok. The horror stories almost always come from people who are working to fulfill needs they do not have and don't understand why their work isn't giving good ROI.
- cupofjoakim 3y agoThe company I worked for when GDPR went into action used kafka in every configuration possible. I've seen some pretty decent uses but also some horrific ones. I agree with the author to an extent - kafka is overly complex for low traffic systems. It can however be a godsend in specific cases, but what those cases are can be hard to pinpoint.
- deleted 3y ago[deleted]
- WaitWaitWha 3y agoFor those who came here for Franz Kafka, this article and posts are about Apache's "distributed event store and stream-processing platform". https://en.wikipedia.org/wiki/Apache_Kafka https://en.wikipedia.org/wiki/Apache_Kafka
- bob88jg 3y agoAmazing reply
- CyberDildonics 3y agoWhy did you think anyone thought this was about Franz Kafka?
- SanderNL 3y ago8000 messages a day, tops? That’s 5 a minute. Does that warrant “infrastructure”? I think a gameboy’s Z80 could handle that load. I don’t want to be dismissive, but I often see these big numbers being posted, like “14M messages” or “thousands of messages” and then adding “per year” or something, which brings it down to toy level load. Even the first “serious” example is about “thousands of messages” per minute. Say 5K a minute. That’s 83 per second, say 100. That seems .. not that interesting? Am I being too dismissive? I think I am. I am not seeing something right. Can anybody say something to widen my perspective?
- wayfinder 3y agoYeah that’s nothing. My busy discussion forum built in PHP running on a toaster of a server was handling way more load than that 15 years ago.
- baz00 3y agoI think you're probably being slightly dismissive. It's not necessarily about load but various other concerns like durability, delivery latency and how failures are handled. There's a big difference between messaging and reliable messaging. I have messaging systems that take 10-20 messages a day but must deliver those messages and do it on a deadline. For that you do need infrastructure (and no that isn't a queue inside a SQL database).
- slt2021 3y agoyou don't need "systems" for 10-20 messages a day. it all could be replaced with S3 buckets and aws-cli with even better durability and delivery latency and error handling than anything you would be able to engineer yourself
- anthonyskipper 3y agoYour point is good, but that stack wouldn't win any latency awards. Many of the people I know using kafka need latencies in the millisecond range.
- jcims 3y agoI've worked at two large companies now with a mature managed Kafka offerings. The 'platform' engineering team handles all of the engineering, implementation, security and compliance, upgrades, observability etc. and have self-service onboarding with lots of recipes and sample integrations. My team moves about 5B messages a day through two topics and we're not putting a dent in the overall volume. It just enables use to move so much more quickly than we would if we had to deal with all of that ourselves. So in our case it's clearly not an anti-pattern, but the right tool for the job.
- browningstreet 3y agoAgreed -- have used it in a bank. It was very suited for that.
- atomicnumber3 3y agoThe fact that you had an entire team operating it is the key, I think. I've seen various big buzzwordy techs used at various shops I've worked at and whether it was nice to work with totally came down to whether there was a team behind it operating it for us. K8s with a k8s team running it - fabulous. Without a team and everyone kind of just needs to know enough to get by, except for the one guy who set it up? Dreadful. Airflow when there's a team running it? Great. Luigi when it's just you, another dude, and the one guy who set it up? Not great. Even RDS is like eh. Still had RDS tip over and we had to do manual vacuuming or something with compacting tables with dead tuples. Annoying when we were paying for RDS to ostensibly not have to do this.
- jpgvm 3y agoAnything is deceptively deep if you understanding never goes beyond skin deep. Also Avro is great but like Kafka you were probably holding it wrong. I do prefer Protobuf in these particular scenarios as Protobufs features more closely align with svc <-> svc RPC style communication patterns while Avro shines in longer lived scenarios where messages need to be archived and you don't want to come up with your own framing for your Protobufs. This is because Avro has the Avro Object Container Format which is a simple block based file format, which allows for relatively efficient seeking, block based compression etc. Protobuf unfortunately doesn't define any standard file formats or even wire protocol framing. If you need to do more than simply store and scan/read in bulk you might want to use Parquet instead though. Reading this blog post was probably a waste of my time, hopefully this comment actually helps someone though.
- safetytrick 3y agoI'm a big Avro fan, it isn't easy to be good at but if you follow its rules it is so much easier to evolve than plain json is.
- tgma 3y agoThis was not explicitly addressed in the post, but the big "Kafka antipattern" out there is building "microservice infrastructure" and using a stateful message broker between services where you should be using RPC/look-aside load balancing with deadlines and retries. Some morons even write books and blog posts about this. The funny thing is this sort of shit is done in the name of scale, but the big folks never operate this way. Large scale infrastructures actively disdain keeping buffers and state in the middle of the request flow. They cannot afford the cost and latency of such systems. They do it the sane way[1]. [1] https://www.usenix.org/conference/osdi23/presentation/saokar https://www.usenix.org/conference/osdi23/presentation/saokar
- waffletower 3y agoCurious what language the OP and their team was using to integrate with Avro. Binary serialization can be a bit awkward, but Avro is a very stable API and it isn't difficult to find/ and or build abstractions to work with it. Perhaps I am spoiled coming from a Clojure perspective?
- tda 3y agoThe main problem I have with Kafka is that their sales team is too good: at my previous employer the CIO was convinced we needed Kafka and bought a contract for sever 100k. But we already had all our events in a postgres database. Admittedly that database had some complicated queries with lots of business logic to get a useful view on the data. But at first I hoped Kafka would somehow make this easier, but of course our particular usecase with a low event volume (hundreds per day), high latency tolerance (next day reporting was considered good enough), highly complex business logic (various computations that required knowledge of what was done previously) all made Kafka just about the least suitable tool for the job. Of course the contract was already signed (I was naturally never consulted up front), so this resulted in lots of solution looking for a problem. No suitable problem was found so I ended up leaving enterprise world for a scale-up and the CIO is still doing whatever he wants for god knows why
- opportune 3y agoOh, I’ve seen much worse than this. I truly believe system design interviews and Confluent marketing/sales have made Kafka a midwit trap: 1. You cannot just use Kafka for free. It will take dev time to set up itself, dev time to code sources and sinks, dev time to handle commonly glossed over but utterly important details like idempotency, retries, duplicate messages, consumed-but-not-committed (or whatever the term is in Kafka world) network interruptions or restarts of consumers. 2. You cannot just continue to use Kafka for free. Running it has an operational cost; this is mitigated using it as a PAAS but not fully solved as you’ll still need to twiddle configurations and scramble to deal with things like “We had no idea we’d need to handle idempotency or use deadletter queues, fixed it, but now need to deal with old data before our fix”. 3. There are many ways you can implement async producer:consumer patterns that are less complex and less costly than Kafka. For example, you can write data to S3. Or you can store records in a regular RDBMS. Kafka isn’t worth it unless you really need “real-time” ingestion but can’t/won’t/shouldn’t implement an actually-real time (synchronous) system instead, like if you get large spikes and are ok with ingestion going from O(seconds) to O(minutes) when that happens. 4. There’s a good chance you don’t need an async queue at all. If consumers can horizontally scale quickly (like with Lambda) why not synchronously invoke them over HTTP/RPC and only use async queues (or a file, etc) for messages that fail multiple retries? Since external users usually aren’t directly writing to your Kafka topic, and thus you have a degree control over your ingestion and consumption, why not just combine the two services (since you can experience data loss from external world to ingestion service anyway, and in fact this is a pretty likely source of failures, your queue may not even be solving the problem you think it is). If you don’t need ordering or partition and are using queues for eg config update propagation, why not just synchronously update consumers or implement basic polling in your consumers? 5. A lot of Kafka/Confluent “features” like retries or logging you can get with so many other tools and services but for some reason these can be the actual selling point more so than the fact it’s an async queue (that also has these features). Yes, in a FAANG design interview where it only costs you 5seconds to say “and we’ll use an async queue like Kafka between these components to handle variable load and partially consumed data” it’s a great tool that saves you a lot of time. And when you pretend integration and maintenance costs don’t exist, and don’t even know what idempotency means, and are sitting across from some slick Confluent salesperson telling you Kafka can be THE database of everything your company does with all these nice features, it sounds great to midwit managers and hasbeen architecture astronauts. In reality? The dumb unsexy alternatives probably solve your actual problem more simply
- deleted 3y ago[deleted]
- FridgeSeal 3y agoWhilst Kafka isn’t a good fit for everything it genuinely sounds like these issues were organisational, not tech and no stack would have suffered the same frustrations.
- tspann 3y agoOn my laptop even a demo with 5000 records a second seems almost too few for Kafka.
- dilyevsky 3y agoKafka will be trash when used as a message bus which is exactly what author had experienced. It can work but it’s designed for stream processing not messaging so it will always be inferior when used this way
- moribvndvs 3y agoThe main struggle I have with Kafka is managing partitioning to avoid hotspots. I have a scenario where we have hundreds of installs of an old and shitty RDBMS on customer sites, and we need to replicate changes to data to a central store. We had to come up with a bespoke event system that would capture insert, update, and delete events, ship them to a REST endpoint, who would then throw them into Kafka to be processed into the central store. Kafka’s ordered message log made it ideal for this scenario, as we can’t play events out of order (although because of poor design in the old databases, it sometimes happens and we built a retry system using additional Kafka topics, nonetheless avoiding out of order messages is critical to keep consumer lag under control). This works mostly ok, but we have a problem when individual customers have big bursts of traffic. Ultimately, we need records to be processed in the order they happen, per customer. Naively, we could partition by customer ID, but arbitrarily adding new partitions as we add customers is not practical over time, and regardless, bulk inserts, updates, etc. could cause large amounts of latency for a customer. So, we’re doing a balancing act of trying to partition using customer ID + a “bundle name” of related tables (the net effect being activity to dependent tables for the same customer always go to the same partition and thus process in order). We’re also looking at using additional topics to create high, medium, and low priority queues, but while that may smooth out some of the problems, it really only breaks the original problem into three smaller versions of the same problem, effectively kicking the can down the road. Ultimately, the best solution would be to get rid of the crappy RDBMS and replace with something that we can binlog or otherwise sync transactionally rather than record by record. We are working on this, but it’s slow going. In the meantime, we continue to wrestle with Kafka partitioning woes. As an aside, we also got rid of Avro. It just didn’t have any benefits that outweighed the challenges to get it and keep it working over time. Much easier to just use plain json, a common message class library between consumers and producers, and a fast, traditional json library. I’ll fully admit that perhaps the avro woes are more an issue inexperience, but I seem to find more people who have the same experience as me than not. Either way, plain json has not caused us any problems.