16 ms·
Why was Apache Kafka created?
- blast 1y ago(I was wondering if this was some sort of generated ripoff but the author worked on Kafka for 6 years: https://x.com/BdKozlovski https://x.com/BdKozlovski.)
- enether 1y agoWhat do you mean by "generated ripoff"? Are you saying it read like AI?
- blast 1y agoI mean that it looked like either AI generated or blogspam or both. Not the writing though, more the look of the page. Happy to be wrong.
- enether 1y agoIt's neither. Only the thumbnail background is AI generated. The look of the page -- that's Substack's default UI, you can't control it too much. The other images are created by me. I'm simply curious what parts give that "cheap" look so I can improve. On Reddit I've had massive amounts of downvotes because they guess the content is AI, when in fact no AI is used in the creation process at all. One guess I have is the bullet points + bolding combo. Most AIs use a ton of that, and rightly so, because it aids in readability.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- polynomial 1y agoWhy was it named that is also a question.
- isaacremuant 1y ago> Jay Kreps chose to name the software after the author Franz Kafka because it is "a system optimized for writing", and he liked Kafka's work. From Wikipedia.
- atombender 1y agoA system optimized for writing could also describe the machine in Kafka's "In the Penal Colony".
- clippy99 1y agoKafkaesque to configure and get running for simple tasks.
- deleted 1y ago[deleted]
- physicles 1y agoThat’s funny, I assumed it was called Kafka because the act of processing items in a queue over and over could be described as kafkaesque.
- theyinwhy 1y agoIt is called Kafka because it can write.
- throw310822 1y agoIt's a bit like calling a dictation software "Hitler" because he also liked to dictate.
- alt227 1y agoThats a brilliant idea, although if I ever create a dictation software I am going to call it 'Mussolini'.
- kachapopopow 1y agoAs someone who has made the mistake of using kafka in a non enterprise space - it really seems like the etcd problem where you need more time to run etcd than to run whatever service you're providing.
- mrweasel 1y agoI previously helped clients setup and run Kafka clusters. Why they'd need Kafka was always our first question, never got a good answer from a single one of them. That's not to say that Kafka isn't useful, it is, in the right setting, but that settings is never "I need a queue". If you need a queue, great, go get RabbitMQ, ZMQ, Redis, SQS, named pipes, pretty anything but Kafka. It's not that Kafka can't do it, but you are making things harder than they needed to be.
- slau 1y agoZMQ is not a managed queue. It’s networking library.
- adev_ 1y ago> Why they'd need Kafka was always our first question, never got a good answer from a single one of them "To follow the hype train, Bro" is often the real answer. > If you need a queue, great, go get RabbitMQ, ZMQ, Redis, SQS, named pipes, pretty anything but Kafka. Or just freaking MQTT. MQTT has been battle-proven for 25 years, is simple and does perfectly the job if you do not ship GBs of blobs through your messaging system (which you should not do anyway).
- atomicnumber3 1y agoIt's resume-driven development. It honestly can make sense for both company and employee. Companies get standard tech stacks people are happy to work with, because working with them gets people experience with tech stacks that are standard at many companies. It's a virtuous cycle. And sure even if you need just a specific thing, it's often better to go slightly overkill for something that's got millions of stack overflow solutions for common issues figured out. Vs picking some niche thing that you are now 1 of like six total people in the entire world using in prod. Obviously the dose makes the poison and don't use kafka for your small internal app thing and don't use k8s where docker will do, but also, probably use k8s if you need more than docker instead of using some weird other thing nobody will know about.
- clippy99 1y agoStartup founder here -- we tried it, and it feels bloated (Java!), bureaucratic and overcomplicated for what it is. Something like Redis queues or even ZMQ probably suffices for 90% of use cases. Maybe in hyper-scaled applications that need to be ultraperformant (e.g., realtime trading, massive streaming platforms) is where Kafka comes into play.
- oulipo2 1y agoHave you tried Redpanda?
- majormajor 1y agoIf you are using this sort of redis queue (https://redis.io/glossary/redis-queue/ https://redis.io/glossary/redis-queue/) with PUSH/POP vs fan-out you're working on a very different sort of problem than what Kafka is built for. Like the article says, fan-out is a key design characteristic. There are "redis streams" now but they didn't exist back then. The durability story and cluster stories aren't as good either, I believe, so they can probably take you so far but won't be as generally suitable depending on where your system goes in the future. There are also things like RedPanda that speak Kafka w/o the Java. However, if you CAN run on a single node w/o worrying about partitioning, you should do that as long as you can get away with it. Once you add multiple partitions ordering becomes hard to reason about and while there are things like message keys to address that, they have limitations and can lead to hotspotting and scaling bottlenecks. But the push/pop based systems also aren't going to give you at-least-once guarantees (looks like Redis at least has a "pop+push" thing to move to a DIFFERENT list that a single consumer would manage but that seems like it gets hairy for scaling out even a little bit...).
- njitbew 1y ago> and it feels bloated (Java!) I'm curious, what exactly feels bloated about Java? I don't feel like the Java language or runtime are particularly bloated, so I'm guessing you're referring to some practices/principles that you often see around Java software?
- slipperydippery 1y ago
- zug_zug 1y agoMy only complaint with this article is that it seems to be implying kafka that linkedIn's problem couldn't have been solved with a bunch of off-the-shelf tools.
- majormajor 1y agoWhat off the shelf tools in 2012 would you propose, exactly?
- tomrod 1y agoSounds like MQTT?
- majormajor 1y agoMQTT wouldn't give you the persistence or the decoupling of fast and slow consumers.
- zug_zug 1y agoMake it less event-orchestrated and use a db. It’s just a social network for recruiters it’s not as complicated as they like to pretend. You don’t need push, it’s just a performance optimization that almost never justifies using a whole new tool.
- AtlasBarfed 1y agoYour solution to a queue and publish subscribe problem is to use a database?
- mrkeen 1y agoAdding onto this. > LinkedIn used site activity data (e.g. someone liked this, someone posted this)1 for many things - tracking fraud/abuse, matching jobs to users, training ML models, basic features of the website (e.g who viewed your profile, the newsfeed), warehouse ingestion for offline analysis/reporting and etc. Who controls the database? Is it the fraud/abuse team responsible for the migrations? Does the ML team tell the Newsfeed team to stop doing so many writes because it's slowing things down?
- dktalks 1y agoI worked on this while I was at LI and I think the major selling point back then was Replayability of messages but it was something similar to what you would get with Pub/Sub. We could also have multiple clients listening and processing same messages for their own purposes so you could use the same queue and have different clients process them as they wanted.
- pojzon 1y agoIts the ability to replay messages at later notice when needed. At least this was the reason we decided to use Kafka instead of simple queues. It was useful when we built new consumer types for the same data we already processed or we knew we gonna have later but cant build now due to prorities.
- varbhat 1y agoDoes anyone use https://nats.io https://nats.io here? I have heard good things about it. I would love to hear about the comparisons between nats.io and kafka
- nchmy 1y agoI dont have kafka experience, but nats is absolutely amazing. Just a complete pleasure to use, in every way. https://www.synadia.com/blog/nats-and-kafka-compared https://www.synadia.com/blog/nats-and-kafka-compared
- dijit 1y agoI got really pissed off with their field CTO for essentially trying to pull the wool over my eyes regarding performance and reliability. Essentially their base product (NATs) has a lot of performance but trades it off for reliability. So they add Jetstream to NATs to get reliability, but use the performance numbers of pure NATs. I got burned by MongoDB for doing this to me, I won’t work with any technology that is marketed in such a disingenuous way again.
- AtlasBarfed 1y agoDon't implement any distributive technology until aphyr has put it through the paces, and even then... Pilot
- munksbeer 1y agohttps://aphyr.com/about https://aphyr.com/about "Unavailable Due to the UK Online Safety Act" :(
- nchmy 1y agoYou mean Jetstream? Can you point to where they are using core NATS numbers to describe Jetstream?
- dijit 1y agoYes, I meant Jetstream (I even typed it but second guessed myself, my mistake) I’m typing these when I get a moment as I’m at a wedding- so I apologise. The issue in the docs was that there are no available Jetstream numbers, so I talked over a video call to the field CTO, who cited the base NATs numbers to me, and when I pressed him on if it was with Jetstream he said that it was without: so I asked for them with Jetstream enabled and he cited the same numbers back to me. Even when I pressed him again that “you just said those numbers are without Jetstream” he said that it was not an issue. So, I got a bit miffed after the call ended, we spent about 45 minutes on the call and this was the main reason to have the call in the first place so I am a bit bent about it. Maybe its better now, this was a year ago.
- BiraIgnacio 1y agoIt was created to teach me the concept of love-hate relationships
- fifilura 1y agoI wanted to write a comment on this topic, but after several tries this thread is where I ended up because it describes my sentiment as well. The arguments in the article are very compelling. But as soon as you choose Kafka you realize the things you hate. Many of the reasons are stupid things - like it uncovers otherwise unimportant bugs in your client code. Or that it just makes experimenting a hassle because it enforces poking around in lots of different places to do something. Or that writing and maintaining the compulsory integration test takes weeks of your time. Sure - you can replay your data - but not until you have fixed all the issues for that special case in your receiving service. I think maybe my main gripe (for us) was that it was a difficult to get an understanding what is actually inside your pipe. Much easier to have that in a solid state in s3? At the end of they day you get annoyed because it slows you down. In particular when you are a small localized team.
- physicles 1y agoTotally agree with this. I’ll add that replaying your data needs special tooling to 1) find the correct offsets on each topic, and 2) spin up whatever daemon will consume that data out-of-band from normal processing, and shut it down when completed. I don’t remember where I read this, but someone made the observation that writing a stream processing system is about 3x harder than writing a batch system, exactly for all the reasons you mentioned. I’m looking at replacing some of our Kafka usage with a clickhouse table that’s ordered and partitioned by insertion time, because if I want to do stuff with that data stream, at least I can do a damn SQL query.
- fifilura 1y agoYes I'll happily extend that to 10x more difficult. At least compared to building a batched pipeline with SQL. I think you should really think hard whether you really need a streaming pipeline. And even if you find that you do, it may be worthwhile to make a batched pipeline as your first implementation. I did exactly what you describe in my previous job. In the beginning with reluctance from our architects who wanted to keep banging the dead horse and did not understand the power of SQL "SQL is not real programming, engineers write java" (ok maybe I deserve a straw-man yellow card here, they don't deserve all of that). But I think they understood after a while. With AWS Athena and Airflow. Good luck, consider me your distant moral support.
- forgetmunch 1y ago[flagged]
- bjourne 1y agoI read it and looked at the block diagrams. I still don't get it. You have "data integration problems". Many software share data. Use database. Problems solved.
- erulabs 1y agoKafka's ability to ingest the firehose and present it as a throttle-able consumable to many different applications is great. If you're thinking "just use a database", it's worth noting that SQL databases are _not well suited_ to drinking from a firehose of writes, and that distributed SQL in 2012 was not a thing. Kafka was one of the first systems that fully embraced the dropping of the C from CAP theorem, which was a big step forward for web applications at scale. If you bristle at that, know that using read-replicas of your postgres database present the same correctness problems. These days though, unless I was at Fortune 100 scale, I'd absolutely turn to Redis Cluster Streams instead. So much simpler to manage and so much cheaper to run. Also I like Kafka because I met two pretty Russian girls in San Francisco a decade back and the group we were in played a game where we described what the company we worked for did in the abstract, and then tried to guess the startup. They said "we write distributed streaming software", I guessed "confluent" immediately. At the time confluent was quite new and small. Fun night. Fun era.
- betaby 1y ago> Kafka's ability to ingest the firehose and present it as a throttle-able consumable to many different applications is great. I sue Kafka precisely for that. Redis Cluster Streams have AOF persistence logs as I see from the doc. How stable it is?
- dxxvi 1y ago> turn to Redis Cluster Streams instead. So much simpler to manage and so much cheaper to run I don't have any experience with Redis Cluster Streams. Could you please tell us how it is simpler to manage? IMO, installing and managing a Kafka cluster in a non Fortune 100 scale is simple enough: run 1 java command for zookeeper, run another java command for a broker (with recent version of Kafka, zookeeper is not needed anymore). The configuration files are not very simple but not very complicated either. When we have another machine, we can run another broker on it. Redis Cluster Streams is cheaper to run because it's written in C, doesn't need a VN to run? Or because its messages are stored in RAM not SSD?
- erulabs 1y ago
- pluto_modadic 1y agolots of people not considering UNPHAT - https://gist.github.com/rponte/67c78e6b3ee349a6e14cb8fb155b7ed1 https://gist.github.com/rponte/67c78e6b3ee349a6e14cb8fb155b7... / https://stevenschwenke.de/pragmaticSoftwareEngineeringUNPHAT https://stevenschwenke.de/pragmaticSoftwareEngineeringUNPHAT - some solutions didn't exist in the past. you evaluate your landscape and your problem, and try to find a good fit.
- chatmasta 1y agoKafka was created at LinkedIn when it was already quite a large web platform, to solve the problem of distributing a unique newsfeed to a wide audience… so while they weren’t Google, they were still closer to Google than the targets of the “UNPHAT” expression. You can apply UNPHAT to most of the people using Kafka today, but it’s not fair to apply it to LinkedIn when they created it.
- pss314 1y agoLinkedIn recently announced that it transitioned from Kafka to Northguard. Introducing Northguard and Xinfra: scalable log storage at LinkedIn [1] & LinkedIn: Stream Processing 4.16.25 [2] [1]: https://www.linkedin.com/blog/engineering/infrastructure/introducing-northguard-and-xinfra https://www.linkedin.com/blog/engineering/infrastructure/int... [2]: https://www.youtube.com/watch?v=RDV6-MUVEbQ https://www.youtube.com/watch?v=RDV6-MUVEbQ
- kafcausir 1y agoI haven't been following kafka for some years now, but i thought Linkedin were heavily invested in it. What happened? Also what happened to Confluent? Their team were ex-Linkedin members from what i remember?
- pm90 1y agoConfluent is expensive and I don’t believe LI used them; they did use OSS kafka. Im guessing that after being acquired by MS they explored other tech.
- gdbsjjdn 1y agoNorthGuard looks like a clean sheet redesign of Kafka. The OSS Kafka community has taken a long time to implement things like KRaft, which addressed metadata scalability concerns by storing metadata in the brokers themselves (it used to be in a separate data store called ZooKeeper which was operationally complicated). NorthGuard also supports splitting and merging ranges of keys without repartitioning the entire existing dataset. The way records are assigned to partitions is a big problem in running Apache Kafka at scale because it requires predicting the key distribution and number of partitions ahead of time. Confluent is 11 years old and IPOed several years ago. It was founded by 3 ex-LinkedIn people who originally designed Kafka. 2 of the 3 founders are still at the company. LinkedIn never used Confluent, Confluent was a company founded to sell an enterprise version of the open source project (and later a cloud version).
- ecoffey 1y agoNorthguard doesn’t look like it’s been open sourced? I’d be curious to know how it compares to Apache Pulsar [0]. I feel like I see some similarities reading the LI blog post. 0: https://pulsar.apache.org/ https://pulsar.apache.org/
- alt227 1y ago> 17 PB/day Are they really generating data at this scale? I cant even imagine a system which creates and stores this much data this fast.
- enether 1y agoWhere does this 17 PB/day number come from? I didn't quote any numbers directly. Looking at the 2012 paper, it implies a 1.35TB per day (they store 9.5TB across all topics at 7d retention)
- djoldman 1y ago> In 2010, LinkedIn had 90 million members. Today, we serve over 1.2 billion members on LinkedIn. Unsurprisingly, this increase has created some challenges over the years, making it difficult to keep up with the rapid growth in the number, volume, and complexity of Kafka use cases. Supporting these use-cases meant running Kafka at a scale of over 32T records/day at 17 PB/day on 400K topics distributed across 10K+ machines within 150 clusters. https://www.linkedin.com/blog/engineering/infrastructure/introducing-northguard-and-xinfra https://www.linkedin.com/blog/engineering/infrastructure/int...
- enether 1y ago~197 GB/s ... nice. I believe these companies save literally every ounce of data they can find. Once you have the infra and teams for it, it seems easy to make a case for storing something. Similarly, Uber has shared they push 89 GB/s through Kafka - 7.7 PB/s. People always ask me - what is a taxi/food-delivery app storing so much
- alt227 1y agoThe biggest capacity HD available today is 30TB. 17PB = about 567 of those drives... being totally filled... per day. I was hoping somebody would come and say this is a simple spelling error or something. The cost of the drives alone seems astronomical, let alone the logistics of the data center keeping up with storing that much data. EDIT: I have just realised that they are probably only processing at this speed, rather than storing it, can anyone confirm if they store all the logs they process?
- game_the0ry 1y agoThis might be hyperbolic, but I think Kafka (or at least the concept of event driven architecture for sharing data across many systems) is one of the most under-rated technologies. It's used at a lot of big corps but is never talked about.
- alphazard 1y agoLike modeling everything as graphs, it is often a trap. Just because a model is flexible enough to capture all of your use cases doesn't mean you should use it. In fact you should prefer less flexible more constrained models that are simpler. Distributed ledgers like Bitcoin do store transitions as events, and that's because nodes need the transitions to valid the next state. So you might say that Bitcoin is a widely run piece of software, using an event driven architecture. Not every system needs to have all of it's state transitions available for for efficient reading. And often times you can derive the state transition from the previous and next state if you really need them (Git does that). Even though Git can compute all of the state transitions for the system, it doesn't store events, it stores snapshots.
- djoldman 1y agoSo now LinkedIn has dropped Kafka and wrote their own called Northguard. More info here: https://www.infoq.com/news/2025/06/linkedin-northguard-xinfra https://www.infoq.com/news/2025/06/linkedin-northguard-xinfr... > According to LinkedIn's engineers, Kafka had become increasingly difficult to manage at LinkedIn's scale (32T records/day, 17 PB/day, 400K topics, 150 clusters). Wtf is LinkedIn doing that they create 17 PB/day?????
- rehathia 1y agoApache Kafka was originally developed by LinkedIn engineers, primarily Jay Kreps, Neha Narkhede, and Jun Rao, around 2010. It was later open-sourced in 2011 and became a top-level project under the Apache Software Foundation in 2012. The creators went on to co-found Confluent, a company that provides commercial support and enterprise features around Kafka.
- dbacar 1y agoThe biggest strength of Kafka in my opinion is consumer groups.I have been using it since 2016 in at least 3 projects and it never failed, not that big workloads though (~100 messages/sec max). However it is a bit difficult to monitor and manage using only the out of the box applications.
- nicoritschel 1y agoI would take a hard look at Kinesis/Pubsub for these data volumes; should cost in the tens of dollars monthly.
- macmar 1y agoIn my experience, Apache Kafka must be understood not as an isolated messaging tool, but as a comprehensive data streaming platform. Its successful implementation demands a holistic approach that encompasses performance, governance, and lifecycle management. I have consistently found that simply adopting the technology without a robust supporting architecture is an ineffective practice that leads to operational challenges. Based on my work managing large-scale Kafka environments across critical sectors, I have identified that their stability and efficiency are upheld by a set of essential practices and tools. These are the non-negotiable pillars for success: Health Checks & Observability: Proactive cluster health monitoring and complete visibility into the data flow are paramount. Failure Management: Implementing dedicated portals and processes for handling Dead-Letter Queues (DLQs) ensures that no critical information is lost during failures. Automation & DevOps: I leverage Strimzi for Kubernetes-native cluster management, orchestrating it through ArgoCD and GitOps practices. This ensures consistent, secure, and repeatable deployments. The correct application of these engineering principles allows for remarkable results. For instance, at a large fashion retail group, I successfully scaled an environment to handle a peak traffic of 480,000 TPS. This high-availability system is efficiently maintained by a lean operational team of just two junior-to-mid-level professionals. From my perspective, success in adopting Kafka is determined by the business context and the maturity of the applied software engineering. The investment in a well-planned architecture and a robust support ecosystem has a clear return, paying for itself through a significant reduction in operational costs (OPEX) within an estimated two-year period. Taming Kafka isn't about new, complex secrets. It's about applying the same robust software engineering and architecture fundamentals we've relied on for +50 years (Software Engineering). The platform is new (2011), the principles are not.
- grantedimmunity 1y ago[dead]
- supriyo-biswas 1y agoThis appears to be a LLM generated comment; if my assumption is correct - please do not do this here. Thank you.