26 ms·
Kafka is Fast – I'll use Postgres
- qsort 11mo agoI feel so seen lol. I work in data engineering and the first paragraph is me all the time. There are a lot of cool technologies (timeseries databases, vector databases, stuff like Synapse on Azure, "lakehouses" etc.) but they are mostly for edge cases. I'm not saying they're useless, but if I see something like that lying around, it's more likely that someone put it there based on vibes rather than an actual engineering need. Postgres is good enough for OpenAI, chances are it's good enough for you.
- zer00eyz 11mo ago> Should You Use Postgres? Most of the time - yes. You should always default to Postgres until the constraints prove you wrong. Kafka, GraphQL... These are the two technology's where my first question is always this: Does the person who championed/lead this project still work here? The answer is almost always "no, they got a new job after we launched". Resume Architecture is a real thing. Meanwhile the people left behind have to deal with a monster...
- kvdveer 11mo agoTo be fair, this is true for all technologically interesting solutions, even when they use postgres. People championing novel solutions typically leave after the window for creativity has closed.
- darkstar_16 11mo agoGraphQL sure, but I'm not sure I'd put kafka in the same bucket. It is a nice technology that has it's use in some cases, where postgresql would not work. It is also something a small team should not start with. Start with postgres and then move on to something else when the need arises.
- forgetfulness 11mo agoWe’re all passing through our jobs, the value of the solutions remains in the hands of the shareholders, if you don’t try to squeeze some long-term value for your resume and long-term employability, you’re assuming a significant opportunity cost on their behalf They’ll be fine if you made something that works, even if it was a bit faddish, make sure you take care of yourself along the way (they won’t)
- candiddevmike 11mo agoAttitudes like this are why management treats developers like children who constantly need to be kept on task, IMO.
- forgetfulness 11mo agoSoftware is a line of work that has astounding amounts of autonomy, if you compare it to working in almost anything else. My point stands, company loyalty tallies up to very little when you’re looking for your next job; no interviewer will care much to hear of how you stood firm, and ignored the siren song of tech and practices that were more modern than the one you were handed down (the tech and practices they’re hiring for). The moment that reverses, I will start advising people not to skill up, as it will look bad in their resumes.
- janwijbrand 11mo ago"resume" as in "resumé" not as in "begin again or continue after a pause or interruption" - it took me longer than I care to admit to get that.
- Groxx 11mo agohaving never hosted a GraphQL service, but I can see many obvious room for problems: is there some reason GraphQL gets so much hate? it always feels to me like it's mostly just a normal RPC system but with some incredibly useful features (pipelining, and super easy to not request data you don't need), with obvious perf issues in code and obvious room for perf abuse because it's easy to allow callers to do N+1 nonsense. so I can see why it's not popular to get stuck with for public APIs unless you have infinite money, it's relatively wide open for abuse, but private seems pretty useful because you can just smack the people abusing it. or is it more due to specific frameworks being frustrating, or stuff like costly parsing and serialization and difficult validation?
- twodave 11mo agoAs someone who works with GraphQL daily, many of the criticisms out there are from before the times of persisted queries, query cost limits, and composite schemas. It’s a very mature and useful technology. I agree with it maybe being less suitable for a public API, but less because of possible abuse and more because simple HTTP is a lot more widely known. It depends on the context, as in all things, of course.
- Groxx 11mo agoyeah, I took one look at it and said "great, so add some cost tracking and kill requests before they exceed it" because like. obviously. it's similar to exposing a SQL endpoint: you need to build for that up front or the obvious results will happen. which I fully understand is more work than "it's super easy just X" which it gets presented as, but that's always the cost of super flexible things. does graphql (or the ecosystem, as that's part of daily life of using it) make that substantially worse somehow? because I've dealt with people using protobuf to avoid graphql, then trying to reimplement parts of its features, and the resulting API is always an utter abomination.
- marcosdumay 11mo agoTake a look on how to implement access control over GraphQL requests. It's useless for anything that isn't public data (at least public for your entire network). And yes, you don't want to use it for public APIs. But if you have private APIs that are so complex that you need a query language, and still want use those over web services, you are very likely doing something really wrong.
- bencyoung 11mo agoKafka is great tech, never sure why people have an issue with it. Would I use it all the time? No, but where it's useful, it's really useful, and opens up whole patterns that are hard to implement other ways
- evantbyrne 11mo agoManaged hosting is expensive to operate and self-managing kafka is a job in of itself. At my last employer they were spending six figures to run three low volume clusters before I did some work to get them off some enterprise features, which halved the cost, but it was still at least 5x the cost of running a mainstream queue. Don't use kafka if you just need queuing.
- CuriouslyC 11mo agoI always push people to start with NATS jetstream unless I 100% know they won't be able to live without Kafka features. It's performant and low ops.
- bencyoung 11mo agoCheapest MSK cluster is $100 a month and can easily run a dev/uat cluster with thousands of messages a second. They go up from there but we've made a lot of use of these and they are pretty useful
- singron 11mo agoI've basically never had a problem with MSK brokers. The issue has usually been "why are we rebalancing?" and "why aren't we consuming?", i.e. client problems.
- evantbyrne 11mo agoIt's not the dev box with zero integrations/storage that's expensive. AWS was quoting us similar numbers for MSK. Part of the issue is that modern kafka has become synonymous with Confluent, and once you buy into those features, it is very difficult to go back. If you're already on AWS and just need queuing, start with SQS.
- sitestable 11mo agoThe best architecture decision is the one that's still maintainable when the person who championed it leaves. Always pretend the person who maintains a project after you knows where you live and all that.
- jjice 11mo agoThis is a well written addition to the list of articles I need to reference on occasion to keep myself from using something new. Postgres really is a startup's best friend most of the time. Building a new product that's going to deal with a good bit of reporting that I began to look at OLAP DBs for, but had hesitation to leave PG for it. This kind of seals it for me (and of course the reference to the class "Just Use Postgres for Everything" post helps) that I should Just Use Postgres (R). On top of being easy to host and already being familiar with it, the resources out there for something like PG are near endless. Plus the team working on it is doing constant good work to make it even more impressive.
- j45 11mo agoIt’s totally reasonable to start with fewer technologies to do more and then outgrow them.
- sanskarix 11mo agoThis mindset is criminally underrated in the startup/indie builder world. There's so much pressure to architect for scale you might never reach, or to use "industry standard" stacks that add enormous complexity. I've been heads-down building a scheduling tool, and the number of times I've had to talk myself out of over-engineering is embarrassing. "Should I use Kafka for event streaming?" No. "Do I need microservices?" Probably not. "Can Postgres handle this?" Almost certainly yes. The real skill is knowing when you've actually outgrown something vs. when you're just pattern-matching what Big Tech does. Most products never get to the scale where these distinctions matter—but they DO die from complexity-induced paralysis. What's been your experience with that inflection point where you actually needed to graduate to more complex tooling? How did you know it was time?
- cpursley 11mo agoRelated: https://www.pgflow.dev https://www.pgflow.dev It's built on pgmq and not married to supabase (nearly everything is in the database). Postgres is enough.
- agentultra 11mo agoYou have to be careful with the approach of using Postgres for everything. The way it locks tables and rows and the serialization levels it guarantees are not immediately obvious to a lot of folks and can become a serious bottle-neck for performance-sensitive workloads. I've been a happy Postgres user for several decades. Postgres can do a lot! But like anything, don't rely on maxims to do your engineering for you.
- sneilan1 11mo agoYes, performance can be a big issue with postgres. And vertical scaling can really put a damper on things when you have a major traffic hit. Using it for kafka is misunderstanding the one of the great uses of kafka which is to help deal with traffic bursts. All of a sudden your postgres server is overwhelmed and the kafka server would be fine.
- zenmac 11mo ago>And vertical scaling can really put a damper on things when you have a major traffic hit. Wouldn't OrioleDB solve that issue though?
- sneilan1 11mo agoNot familiar with OrioleDB. I’ll look it up. May I ask how this helps? Just curious.
- mike_hearn 11mo agoIt's worth noting that Oracle has solved this problem. It has horizontal multi-master scalability (not sharded) and a queue subsystem called TxEQ which scales like Kafka does, but it's also got the features of a normal MQ broker. You can dequeue a message into a transaction, update tables in that same transaction, then commit to remove the message from the queue permanently. You can dequeue by predicate, delay messages, use producer/consumer patterns etc. It's quite flexible. The queues can be accessed via SQL stored procs, or client driver APIs, or it implements a Kafka compatible API now too I think. If you rent a cloud DB then it can scale elastically which can make this cheaper than Postgres, believe it or not. Cloud databases are sold at the price the market will bear not the cost of inputs+margin, so you can end up paying for Postgres as much as you would for an Oracle DB whilst getting far fewer features and less scalability. Source: recently joined the DB team at Oracle, was surprised to learn how much it can do.
- guywithahat 11mo ago> One camp chases buzzwords > ... > The other camp chases common sense I don't really like these simplifications. Like one group obviously isn't just dumb, they're doing things for reasons you maybe don't understand. I don't know enough about data science to make a call, but I'm guessing there were reasons to use Kafka due to current hardware limits or scalability concerns, and while the issues may not be as present today that doesn't mean they used Kafka just because they heard a new word and wanted to repeat it.
- sumtechguy 11mo agoKafka and other message systems like it have their uses. But sometimes all you need is just need a database. Now you start doing realtime streaming and notifications and event type things a messaging system is good. You can even back it up with a boring database. Would I start with kafka? Probably not. I would start with a boring databsee and then if if my bashing on the db over and over saying 'have you changed' doesnt work as good anymore then you put in a messaging system.
- temporallobe 11mo agoAgree with this sentiment - it’s easy to be judgmental about these things, but project-level issues and decisions can be very complicated and engineers often have little to no visibility into them. We’re using Kafka for a gigantic pipeline where IMO any reasonably modern database would suffice (and may even be superior), but our performance requirements are unclear. At some point in the distant future, we may have a significant surge in data quantity and speed, requiring greater throughput and (de)serialization speed, but I am not convinced that Kafka ultimately helps us there. I imagine this is a case where the program leadership was sold a solution which we are now obligated to use. This happens a LOT, and I have seen unnecessary and unused products cost companies millions over the years. For example, my team was doing analysis on replacing our existing Atlassian Data Center with other solutions, and in doing so, we discovered several underused/unused Atlassian plugins for which we are paying very high license fees. At some point, users over the years had requested some functionality for a specific workflow and the plugins were purchased. The people and projects went away or otherwise processes became OBE, but the plugins happily hummed along while the bills were paid.
- sneilan1 11mo agoI'm starting to like mongodb a lot more given the python library mongomock. I find it wonderful to create tests that run my queries against mongo in code before I deploy them. Yes, mongo has a lot of quirks and you have to know aws networking to set it up with your vpc so you don't get nailed with egress costs. And it's not the same query patterns and some queries are harder and you have maintain your own schemas. But the ability to test mongo code with mongomock w/o having to run your own mongo server is SO VALUABLE. And yes, there are edge cases with mongomock not supporting something but the library is open source and pretty easy to modify. And it fails loudly which is super helpful. So if something is not supported you'll know. Maybe you might find a real nasty feature that's hard to implement but then just use a repository pattern like you would for testing postgres code in your application. https://github.com/mongomock/mongomock https://github.com/mongomock/mongomock Extrapolating from my personal usage of this library to others, I'm starting to think that mongodb's 25 billion dollar valuation is partially based on this open source package :)
- pphysch 11mo agoOr just use devcontainers and have an actual Postgres DB to test against? I've even done this on a Chromebook. This is a solved problem.
- sneilan1 11mo agoTrue but then my tests take longer to run. I really like having very fast tests. And then my tests have to make local network calls to a postgres server. I like my tests isolated.
- pphysch 11mo agoThey are isolated, your devcontainer config can live in your source repo. And you're not gonna see significant latency from your loopback interface... If your test suite includes billions of queries you may want to reassess.
- 11mo ago
- honkostani 11mo agoResume driven design, is running into the desert of moores plateau punishing the use of ever more useless abstractions. They get quieter, because their projects keep on dying after the revolutionary tech is introduced and they jump ship.
- deleted 11mo ago[deleted]
- jimbokun 11mo agoFor me the killer feature of Kafka was the ability to set the offset independently for each consumer. In my company most of our topics need to be consumed by more than one application/team, so this feature is a must have. Also, the ability to move the offset backwards or forwards programmatically has been a life saver many times. Does Postgres support this functionality for their queues?
- Jupe 11mo agoIsn't it just a matter of having each consumer use their own offset? I mean if the queue table is sequentially or time-indexed, the consumer just provides a smaller/earlier key to accomplish the offset? (Maybe I'm missing something here?)
- altcognito 11mo agoCorrect, offsets and sharding aren't magic. And partitions in Kafka are user defined, just like they would be for postgresql.
- jimbokun 11mo agoYes. Is a queuing system baked into Postgres? Or there client libraries that make it look like one? And do these abstractions allow for arbitrarily moving the offset for each consumer independently? If you're writing your own queuing system using pg for persistence obviously you can architect it however you want.
- cortesoft 11mo agoKafka allows you to have a consumer group… you can have multiple workers processing messages in parallel, and if they all use the same group id, the messages will be sharded across all the workers using that key… so each message will only be handled by one worker using that key, and every message will be given to exactly one worker (with all the usual caveats of guaranteed-processed-exactly-once queues). Other consumers can use different group keys and they will also get every single message exactly once. So if you want an individual offset, then yes, the consumer could just maintain their own… however, if you want a group’s offset, you have to do something else.
- deleted 11mo ago[deleted]
- johnyzee 11mo agoSeems like you would at the very least need a fairly thick application layer on top of Postgres to make it look and act like a messaging system. At that point, seems like you have just built another messaging system. Unless you're a five man shop where everybody just agrees to use that one table, make sure to manage transactions right, cron job retention, YOLO clustering, etc. etc. Performance is probably last on the list of reasons to choose Kafka over Postgres.
- j45 11mo agoYou expose the api on Postgres much like any other group of developers use and call it a day. There’s several implementations of queues to increase the chance of finishing what one is after. https://github.com/dhamaniasad/awesome-postgres https://github.com/dhamaniasad/awesome-postgres
- dagss 11mo agoThere's a lot of logic involved client side regarding managing read cursors and marking events as processed consumer side. Possibly also client side error queues and so on. I truly miss a good standard client side library following the Kafka-in-SQL philosophy. I started on in my previous job and we used it internally but it never got good enough that it would be widely used elsewhere, and now I work somewhere else... (PS: Talking about the pub/sub Kafka-like usecase, not the work queue FOR UPDATE usecase)
- odie5533 11mo agoHow fast is failover?
- vbezhenar 11mo agoHow do you implement "unique monotonically-increasing offset number"? Naive approach with sequence (or serial type which uses sequence automatically) does not work. Transaction "one" gets number "123", transaction "two" gets number "124". Transaction "two" commits, now table contains "122", "124" rows and readers can start to process it. Then transaction "one" commits with its "123" number, but readers already past "124". And transaction "one" might never commit for various reasons (e.g. client just got power cut), so just waiting for "123" forever does not cut it. Notifications can help with this approach, but then you can't restart old readers (and you don't need monotonic numbers at all).
- theK 11mo ago> unique monotonically-increasing offset number Isn't it a bit of a white whale thing that a umion can solve all one's subscriber problems? Afaik even with kafka this isn't completely watertight.
- sigseg1v 11mo agoWhat about a `DEFERRABLE INITIALLY DEFERRED` trigger that increments a sequence only on commit?
- singron 11mo agoThe log_counter table tracks this. It's true that a naive solution using sequences does not work for exactly the reason you say.
- xnorswap 11mo agoIt's a tricky problem, I'd recommend reading DDIA, it covers this extensively: https://www.oreilly.com/library/view/designing-data-intensive-applications/9781491903063/ https://www.oreilly.com/library/view/designing-data-intensiv... You can generate distributed monotonic number sequences with a Lamport Clock. https://en.wikipedia.org/wiki/Lamport_timestamp https://en.wikipedia.org/wiki/Lamport_timestamp The wikipedia entry doesn't describe it as well as that book does. It's not the end of the puzzle for distributed systems, but it gets you a long way there. See also Vector clocks. https://en.wikipedia.org/wiki/Vector_clock https://en.wikipedia.org/wiki/Vector_clock Edit: I've found these slides, which are a good primer for solving the issue, page 70 onwards "logical time": https://ia904606.us.archive.org/32/items/distributed-systems-dr-martin-kleppmann/dist-sys-handout.pdf https://ia904606.us.archive.org/32/items/distributed-systems...
- ownagefool 11mo agoThe camps are wrong. There's poles. 1. Is folks constantly adopting the new tech, whatever the motivation, and 2. I learned a thing and shall never learn anything else, ever. Of course nobody exists actually on either pole, but the closer you are to either, the less pragmatic you are likely to be.
- wosined 11mo agoI am the third pole: 3. Everything we have currently sucks and what is new will suck for some hitherto unknown reason.
- ownagefool 11mo agoHeh, me too. I think it's still just 2 poles. However, I probably shouldn't have prescribed motivation to latter pole, as I purposely did not with the former. Pole 2 is simply never adopt anything new ever, for whatever the motivation.
- antonvs 11mo agoIf you choose wisely, things should suck less overall as you move forward. That's kind of the overall goal, otherwise we'd all still be toggling raw machine code into machines using switches.
- wosined 11mo agoComputers got faster, software is not so straightforward. I don't even know why a text webpage needs 100mb of memory to render and display.
- binarymax 11mo agoThis is it right here. My foil is the Elasticsearch replacement because PG has inverted indices. The ergonomics and tunability of these in PG are terrible compared to ES. Yes, it will search, but I wouldn’t want to be involved in constructing or maintaining that search.
- jppope 11mo ago
- uberduper 11mo agoHas this person actually benchmarked kafka? The results they get with their 96 vcpu setup could be achieved with kafka on the 4 vcpu setup. Their results with PG are absurdly slow. If you don't need what kafka offers, don't use it. But don't pretend you're on to something with your custom 5k msg/s PG setup.
- loire280 11mo agoIn fact, a properly-configured Kafka cluster on minimal hardware will saturate its network link before it hits CPU or disk bottlenecks.
- j45 11mo agoBut it can do so many processes a second I’ll be able to scale to the moon before I ever launch.
- theK 11mo agoIsn't that true for everything on the cloud? I thought we are long into the era where your disk comes over the network there.
- UltraSane 11mo agoA network link can be anything from 1Gbps to 800Gbps.
- altcognito 11mo agoThis doesn't even make sense. How do you know what the network links or the other bottlenecks are like? There are a grandiose number of assumptions being made here.
- loire280 11mo agoThere is a finite and relatively narrow range of ratios of CPU, memory, and network throughput in both modern cloud offerings and bare hardware configurations. Obviously it's possible to build, for example, a machine with 2 cores, a 10Gbps network link, and a single HDD that would falsify my statement.
- rudderdev 11mo agoDiscussion on the same topic "Postgres over Kafka" - https://news.ycombinator.com/item?id=44445841 https://news.ycombinator.com/item?id=44445841
- dangoodmanUT 11mo ago96 cores to get 240MB/s is terrible. Redpanda can do this with like one or two cores
- greenavocado 11mo agoRedpanda might be good (I don't know) but I threw up a little in my mouth when I opened their website and saw "Build the Agentic Data Plane"
- School-Cotton 11mo agoThe marketing website of every data-related startup sounds like that now. I agree it’s dumb, but you can safely ignore it.
- enether 11mo agohehe, yeah it is. I could have probably got a GB/s out of that if I ran it properly - but it's at the scale where you expect it to be terrible due to the mismatch of workloads
- loftsy 11mo agoI am about to start a project. I know I want an event sourced architecture. That is, the system is designed around a queue, all actors push/pull into the queue. This article gives me some pause. Performance isn't a big deal for me. I had assumed that Kafka would give me things like decoupling, retry, dead-lettering, logging, schema validation, schema versioning, exactly once processing. I like Postgres, and obviously I can write a queue ontop of it, but it seems like quite a lot of effort?
- j45 11mo agoIt might look like a lot of effort, but if you follow a tutorial/YouTube video step by step you will be surprised. It’s mostly registering the Postgres database functions which is one time. There are also pre-made Postgres extensions that already run the queue. These days i would like consider m starting with Supabase self hosted which has the Postgres ready to tweak.
- mkozlows 11mo agoIf what you want is a queue, Kafka might be overkill for your needs. It's a great tool, but it definitely has a lot of complexity relative to a straightforward queue system.
- mrkeen 11mo agoEvent-sourcing != queue. Event-sourcing is when you buy something and get a receipt, you go stick it in a shoe-box for tax time. A queue is you get given receipts, and you look at them in the correct order before throwing each one away.
- loftsy 11mo agoTrue. I think my system is sort of both. I want to put some events in a queue for a finite set of time, process them as a single consolidated set, and then drop them all from the queue.
- singron 11mo agoKafka also doesn't give you all those things. E.g. there is no automatic dead-lettering, so a consumer that throws an exception will endlessly retry and block all progress on that partition. Kafka only stores bytes, so schema is up to you. Exactly-once is good, but there are some caveats (you have to use kafka transactions, which are significantly different than normal operation, and any external system may observe at-least-once semantics instead). Similar exactly-once semantics would also be trivial in an RDBMS (i.e. produce and consume in same transaction). If you plan on retaining your topics indefinitely, schema evolution can become painful since you can't update existing records. Changing the number of partitions in a topic is also painful, and choosing the number initially is a difficult choice. You might want to build your own infrastructure for rewriting a topic and directing new writes to the new topic without duplication. Kafka isn't really a replacement for a database or anything high-level like a ledger. It's really a replicated log, which is a low-level primitive that will take significant work to build into something else.
- CuriouslyC 11mo agoIf you don't need all the bells and whistles of Kafka, NATS Jetstream is usually the way to go.
- misja111 11mo ago> One camp chases buzzwords .. the other common sense How is it common sense to try to re-implement Kafka in Posgres? You probably need something similar but more simple. Then implement that! But if you really need something like Kafka, then .. use Kafka! IMO the author is now making the same mistake as some Kafka evangelists that try to implement a database in Kafka.
- enether 11mo agoI’m making the example of a pub sub system. I’m most familiar with Kafka so drew parallels to it. I didn’t actually implement everything Kafka offers - just two simple pub sub like queries.
- shikhar 11mo agoPostgres is a way better fit than Kafka if you want a large number of durable streams. But a flexible OLTP database like PG is bound to require more resources and polling loops (not even long poll!) are not a great answer for following live updates. Plug: If you need granular, durable streams in a serverless context, check out s2.dev
- dagss 11mo agos2.dev looks cool... I jumped around the home page a bit and couldn't perfectly grasp what it is quickly though. But if it is about decoupling the Kafka approach and client side libraries from the use of Kafka specifically I am cheering for you. Could you see using the s2.dev protocol on top of services using SQL in the way of the article, assigning event sequence numbers, as a good fit? Or is s2 fundamentally the component that assigns event numbers? I feel like we tried to do something similar to you, but for SQL DBs, but am not sure: https://github.com/vippsas/feedapi-spec https://github.com/vippsas/feedapi-spec
- this_user 11mo agoThe real two camps seem to be: 1) People constantly chasing the latest technology with no regard for whether it's appropriate for the situation. 2) People constantly trying to shoehorn their favourite technology into everything with no regard for whether it's appropriate for the situation.
- j45 11mo agoKafka is anything but new. It does get shoehorned too. Postgres also has been around for a long time and a lot of people didn’t know all it can do which isn’t what we normally think about with a database. Appropriateness is a nice way to look at it as long as it’s clear whether or not it’s about personal preferences and interpretations and being righteous towards others with them. Customers rarely care about the backend or what it’s developed in, except maybe for developer products. It’s a great way to waste time though.
- PeterCorless 11mo ago2) above is basically "Give a kid a hammer, and everything becomes a nail." The third camp: 3) People who look at a task, then apply a tool appropriate for the task.
- me551ah 11mo agoImagine if historic humans had decided that only hammers are enough. That there is no need for a specialized tool like Scissors, Chisel, Axe, Wrench, Shovel , Sickle and that a hammer and fingers are enough. Use the tool which is appropriate for the job, it is trivial to write code to use them with LLMs these days and these software are mature enough to rarely cause problems and tools built for a purpose will always be more performant.
- justinhj 11mo agoAs engineers we should try to use the right tool for the job, which means thinking about the development team's strengths and weaknesses as well as differentiating factors your product should focus on. Often we are working in the cloud and it's much easier to use a queue or a log database service than manage a bunch of sql servers and custom logic. It can be more cost effective too once you factor in the development time and operational costs. The fact that there is no common library that implements the authors strategy is a good sign that there is not much demand for this.
- dzonga 11mo agowhat's not spoken about in the above article ? ease of use. in ruby If I want to use kafka I can use karafka. or redis streams via the redis library. likewise if kafka is too complex to run there's countless alternatives which work as well - hell even 0mq with client libraries. now with the postgres version I have to write my own stuff which I might not where it's gonna lead me. postgres is scalable, no one doubts that. but what people forget to mention is the ecosystem around certain tools.
- j45 11mo agoI’m not sure where it says you have to write your own stuff, there seem to be some of queues with libraries. https://github.com/dhamaniasad/awesome-postgres https://github.com/dhamaniasad/awesome-postgres There is at least a Python example here.
- dagss 11mo agoWork queues is easy. It is significantly more work for the client side implementation of event log consumers, which the article also talk about. For instance persisting the client side cursors. And I have not seen widely used standard implementations of those. (I started one myself once but didn't finish.)
- enether 11mo agoThat's true. There seems to be two planes of ease of use - the app layer (library) and the infra layer (hosting). The app layer for Postgres is still in development, so if you currently want to run pub-sub (Kafka) on it, it will be extra work to develop that abstraction. I hope somebody creates such a library. It's a one-time cost but then will make it easier for everybody.
- losvedir 11mo agoMaybe I missed it in the design here, but this pseudo-Kafka Postgres implementation doesn't really handle consumer groups very well. The great thing about Kafka consumer groups is it makes it easy to spread the load over several instances running your service. They'll all connect using the same group, and different partitions will be assigned to the different instances. As you scale up or down, the partition responsibilities will be updated accordingly. You need some sort of server-side logic to manage that, and the consumer heartbeats, and generation tracking, to make sure that only the "correct" instances can actually commit the new offsets. Distributed systems are hard, and Kafka goes through a lot of trouble to ensure that you don't fail to process a message.
- mrkeen 11mo agoRight, the author's worldview is that Kafka is resume-driven development, used by people "for speed" (even though they are only pushing 500KB/s). Of course the implementation based off that is going to miss a bit.
- jasonthorsness 11mo agoUsing a single DBMS for many purposes because it is so flexible and “already there” from an operations perspective is something I’ve seen over and over again. It usually goes wrong eventually with one workload/use screwing up others but maybe that’s fine and a normal part of scaling? I think a bigger issue is the DBMS themselves getting feature after feature and becoming bloated and unfocused. Add the thing to Postgres because it is convenient! At least Postgres has a decent plugin approach. But I think more use cases might be served by standalone products than by add-ons.
- quaunaut 11mo agoIt's a normal part of scaling because often bringing in the new technology introduces its own ways of causing the exact same problems. Often they're difficult to integrate into automated tests so folks mock them out, leading to issues. Or a configuration difference between prod/local introduces a problem. Your DB on the other hand is usually a well-understood part of your system, and while scaling issues like that can cause problems, they're often fairly easy to predict- just unfortunate on timing. This means that while they'll disrupt, they're usually solved quickly, which you can't always say for additional systems.
- munchbunny 11mo agoMy general opinion, off the cuff, from having worked at both small (hundreds of events per hour) and large (trillions of events per hour) scales for these sorts of problems: 1. Do you really need a queue? (Alternative: periodic polling of a DB) 2. What's your event volume and can it fit on one node for the foreseeable future, or even serverless compute (if not too expensive)? (Alternative: lightweight single-process web service, or several instances, on one node.) 3. If it can't fit on one node, do you really need a distributed queue? (Alternative: good ol' load balancing and REST API's, maybe with async semantics and retry semantics) 4. If you really do need a distributed queue, then you may as well use a distributed queue, such as Kafka. Even if you take on the complexity of managing a Kafka cluster, the programming and performance semantics are simpler to reason about than trying to shoehorn a distributed queue onto a SQL DB.
- oulipo2 11mo agoI want to rewrite some of my setup, we're doing IoT, and I was planning on MQTT -> Redpanda (for message logs and replay, etc) -> Postgres/Timescaledb (for data) + S3 (for archive) (and possibly Flink/RisingWave/Arroyo somewhere in order to do some alerting/incrementally updated materialized views/ etc) this seems "simple enough" (but I don't have any experience with Redpanda) but is indeed one more moving part compared to MQTT -> Postgres (as a queue) -> Postgres/Timescaledb + S3 Questions: 1. my "fear" would be that if I use the same Postgres for the queue and for my business database, the "message ingestion" part could block the "business" part sometimes (locks, etc)? Also perhaps when I want to update the schema of my database and not "stop" the inflow of messages, not sure if this would be easy? 2. also that since it would write messages in the queue and then delete them, there would be a lot of GC/Vacuuming to do, compared to my business database which is mostly append-only? 3. and if I split the "Postgres queue" from "Postgres database" as two different processes, of course I have "one less tech to learn", but I still have to get used to pgmq, integrate it, etc, is that really much easier than adding Redpanda? 4. I guess most Postgres queues are also "simple" and don't provide "fanout" for multiple things (eg I want to take one of my IoT message, clean it up, store it in my timescaledb, and also archive it to S3, and also run an alert detector on it, etc) What would be the recommendation?
- ayongpm 11mo agoJust dropping this here casually: sup { position: relative; top: -0.4em; line-height: 0; vertical-align: baseline; }
- oulipo2 11mo agoI want to rewrite some of my setup, we're doing IoT, and I was planning on MQTT -> Redpanda (for message logs and replay, etc) -> Postgres/Timescaledb (for data) + S3 (for archive) (and possibly Flink/RisingWave/Arroyo somewhere in order to do some alerting/incrementally updated materialized views/ etc) this seems "simple enough" (but I don't have any experience with Redpanda) but is indeed one more moving part compared to MQTT -> Postgres (as a queue) -> Postgres/Timescaledb + S3 Questions: 1. my "fear" would be that if I use the same Postgres for the queue and for my business database, the "message ingestion" part could block the "business" part sometimes (locks, etc)? Also perhaps when I want to update the schema of my database and not "stop" the inflow of messages, not sure if this would be easy? 2. also that since it would write messages in the queue and then delete them, there would be a lot of GC/Vacuuming to do, compared to my business database which is mostly append-only? 3. and if I split the "Postgres queue" from "Postgres database" as two different processes, of course I have "one less tech to learn", but I still have to get used to pgmq, integrate it, etc, is that really much easier than adding Redpanda? 4. I guess most Postgres queues are also "simple" and don't provide "fanout" for multiple things (eg I want to take one of my IoT message, clean it up, store it in my timescaledb, and also archive it to S3, and also run an alert detector on it, etc) What would be the recommendation?
- Copenjin 11mo agoI'm not really convinced by the comment on NOTIFY instead of the inferior (at least in theory) polling, I expect the global queue if it's really global to be only a temporary location to collect notifications before sending them and not a bottleneck. Never did any benchmark with PG or Oracle (that has a similar feature) but I expect that depending on the polling frequency and average amount of updates each solution could be the best depending on the circumstances.
- sc68cal 11mo ago> Postgres doesn’t seem to have any popular libraries for pub-sub9 use cases, so I had to write my own. Ok so instead of running Kafka, we're going to spend development cycles building our own?
- enether 11mo agoIt would be nice if a library like pgmq got built. Not sure what the demand for that is, but it feels like there may be a niche
- 8cvor6j844qw_d6 11mo ago> Should You Use Postgres? > Most of the time - yes. You should always default to Postgres until the constraints prove you wrong. Interesting. I've also been by my seniors that I should go with PostgreSQL by default unless I have a good justification not to.
- heyitsdaad 11mo agoIf the only tool you know is a hammer, everything starts looking like a nail.
- bleonard 11mo agoI am excited about the Rails defaults where background and cache and sockets are all database driven. For normal-sized projects that still need those things, it's a huge win in simplicity.
- psadri 11mo agoA resource that would benefit the entire community is a set of ballpark figures for what kind of performance is "normal" given a particular hardware + data volume. I know this is a hard problem because there is so much variation across workloads, but I think even order of magnitude ballparks would be useful. For example, it could say things like: task: msg queue software: kafka hardware: m7i.xlarge (vCPUs: 4 Memory: 16 GiB) payload: 2kb / msg possible performance: ### - #### msgs / second etc… So many times I've found myself wondering: is this thing behaving within an order of magnitude of a correctly setup version so that I can decide whether I should leave it alone or spend more time on it.
- phendrenad2 11mo agoSince everyone is offering what they think the "camps" should be, here's another perspective. There are two camps: (A) Those who look at performance metrics ("96 cores to get 240MB/s is terrible") and assume that performance itself is enough to justify overruling any other concern (B) Those who look at all of the tradeoffs, including budget, maintenance, ease-of-use, etc. You see this a lot in the tech world. "Why would you use Python, Python is slow" (objectively true, but does it matter for your high-value SaaS that gets 20 logins per day?)
- wagwang 11mo agoIsn't listen/notify absurdly slow and lock contentious
- ryandvm 11mo agoI think my only complaint about Kafka is the widespread misunderstanding that it is a suitable replacement for a work queue. I should not be having to explain to an enterprise architect the distinction between a distributed work queue and event streaming platform.
- lisbbb 11mo agoIt's not so much that they don't know as it they think Kafka is sexier, or, in my case, it was mandated to use it for everything because they were paying for the cluster. I solved one problem, very flexibly, in Elastic and they weren't even interested at all. It was Kafka or nothing. That's reality in a lot of companies.
- deleted 11mo ago[deleted]
- Sparkyte 11mo agoYou can also use Redis as a queue if the data isn't in danger of being too important.
- jdboyd 11mo agoWhile I appreciate the Postgres for everything point of view, and most of the times I use other things it could fit in Postgres, there are two areas that keep me using RabbitMQ, Redis, or a something like Elastic. First, I frequently use Celery and Celery doesn't support using Postgres as a broker. It seems like it should, but I guess no one has stepped up to write that. So, when I use Celery, I end up also using Redis or RabbitMQ. Second, if I need mqtt clients coming in from the internet at large, I don't feel comfortable exposing Postgres to that. Also, I'd rather use the mqtt ecosystem of libraries rather than having all of those devices talk Postgres directly. Third, sometimes I want a size constrained memory only database or a database that automatically expires untouched records, and for either of those I usually use Redis. For these two tasks I use Redis. I imagine that it would be worth making a reusable set of stored procedures to accomplish the auto-expiring of unused records, but I haven't implemented it. I have no idea how to make Postgres be memory memory only with a constrained memory side.
- nyrikki 11mo ago> The claim isn’t that Postgres is functionally equivalent to any of these specialized systems. The claim is that it handles 80%+ of their use cases with 20% of the development effort. (Pareto Principle) Lots of us that built systems when SQL was the only option, know that doesn’t hold overtime. SStable backed systems have their applications, and I have never seen dedicated Kafka teams like we used to have with DBAs We have the tools to make decisions based on real tradeoffs. I highly recommend people dig into the appropriate tools to select vs making pre-selected products fit an unknown problem domain. Tools are tactics, not strategies, tactics should be changeable with the strategic needs.
- mbo 11mo agoThis is an article in desperate need for some data visualizations. I do not think it does an effective job of communicating differences in performance.
- lisbbb 11mo agoIf you are doing high volume, there is no way that a SQL db is going to keep up. I did a lot of work with Kafka but what we constantly ran into was managing expectations--costs were higher, so the business needs to strongly justify why they need their big data toy, and joins are much harder, as well as data validation in real time. It made for a frustrating experience most of the time--not due to the tech as much as dealing with people who don't understand the costs and benefits. On the major projects I worked on, we were "instructed" to use Kafka for, I guess, internal political reasons. They already had Hadoop solutions that more or less worked, but the code was written by idiots in "Spark/Scala" (their favorite buzzword to act all high and mighty) and that code had zero tests (it was truly a "test in prod" situation there). The Hadoop system was managed by people who would parcel out compute resources politically, as in, their friends got all they wanted while everyone else got basically none. This was a major S&P company, Fortune 10, and the internal politics were abusive to say the least.
- rjurney 11mo agoOne bad message in a Kafka queue and guess what? The entire queue is down because it kills your workers over and over. To fix it? You have to resize the queue to zero, which means losing requests. This KILLS me. Jay Kreps says there is no reason it can't be fixed, but it never had been and this infuriates me because it happens so often :)
- pram 11mo agoYou can modify a consumer groups offset to any value JFYI, so you really don’t need to purge the topic. You can just start after the bad message.
- deleted 11mo ago[deleted]
- jackvanlightly 11mo ago> A 500 KB/s workload should not use Kafka This is a simplistic take. Kafka isn't just about scale, it, like other messaging systems provide queue/streaming semantics for applications. Sure you can roll your own queue on a database for small use cases, but it adds complexity to the lives of developers. You can offload the burden of running Kafka by choosing a Kafka-as-a-service vendor, but you can't offload the additional work of the developer that comes from using a database as a queue.
- cyanf 11mo agoThere are existing solutions for queues in Postgres, notably pgmq.
- enether 11mo agoThe question is the organizational overhead in adopting yet another specialized distributed system, which btw frequently is about scalability at its core. Kafka's original paper emphasizes this ("We introduce Kafka, a distributed messaging system that we developed for collecting and delivering high volumes of log data with low latency. ", "We made quite a few unconventional yet practical design choices in Kafka to make our system efficient and scalable.")[1] To be honest, there isn't a large burden in running Kafka when it's 500 KB/s. The system is so underutilized there's nothing to cause issues with it. But regardless, the organizational burden persists. As the piece mentions - "Managed SaaS offerings trade off some of the organizational overhead for greater financial costs - but they still don’t remove it all.". Some of the burden continues to exist even if a vendor hosts the servers for you. The API needs to be adopted, the clients have many configs, concepts like consumer groups need to be understood, the vendor has its own UI, etc. The Kafka API isn't exactly the simplest. I wouldn't recommend people write the pub-sub-on-postgres SQL themselves - a library should abstract it away. What is the complexity being added from a library with a simple API? Regardless if that library is based on top of Postgres, Kafka or another system - precisely what complexity is added to the lives of developers? I really don't see any complexity existing at this miniscule scale, neither at the app developer layer or the infra operator layer. But of course, I haven't run this in production so I could be wrong. [1] - https://notes.stephenholiday.com/Kafka.pdf https://notes.stephenholiday.com/Kafka.pdf
- deleted 11mo ago[deleted]
- jeeybee 11mo agoIf you like the “use Postgres until it breaks” approach, there’s a middle ground between hand-rolling and running Kafka/Redis/Rabbit: PGQueuer. PGQueuer is a small Python library that turns Postgres into a durable job queue using the same primitives discussed here — `FOR UPDATE SKIP LOCKED` for safe concurrent dequeue and `LISTEN/NOTIFY` to wake workers without tight polling. It’s for background jobs (not a Kafka replacement), and it shines when your app already depends on Postgres. Nice-to-haves without extra infra: per-entrypoint concurrency limits, retries/backoff, scheduling (cron-like), graceful shutdown, simple CLI install/migrations. If/when you truly outgrow it, you can move to Kafka with a clearer picture of your needs. Repo: https://github.com/janbjorge/pgqueuer https://github.com/janbjorge/pgqueuer Disclosure: I maintain PGQueuer.
- bmcahren 11mo agoA huge benefit of single-database operations at scale is point-in-time recovery for the entire system thereby not having to coordinate recovery points between data stores. Alternatively, you can treat your queue as volatile depending on the purpose.
- nchmy 11mo agoSeems like instead of a hand-rolled, polling Pub/sub, could instead do CDC instead with a golang logical replication/cdc library. There's surely various. Or just use NATS for queues and pubsub - dead simple, can embed in your Go app and does much more than Kafka
- brikym 11mo agoIf you don't mind Redis then use Redis Streams. It gives you an eventlog without worrying about postgres performance issues and has consumer groups.
- tele_ski 11mo agoBeen using valkey streams recently and loving it. Took a bit to understand how to to properly use it but now that I've figured it out I'd highly recommend trying it. It's very easy to setup and get going and just works.
- tarun_anand 11mo agoCouldn't agree more. Have built and ran an in-house postgresql based queue for several years. It can handle 5-10k msg/s in our production workloads.
- smoyer 11mo agoKafka is fast ... And MongoDB is web scale [0]. I completely agree that we shouldn't go chasing each new technical bauble but we are also wasting breath on those that do. 0. https://youtu.be/b2F-DItXtZs?si=vrB-UxCHIgMYGKFt https://youtu.be/b2F-DItXtZs?si=vrB-UxCHIgMYGKFt
- spectraldrift 11mo ago> Should You Use Postgres? Most of the time - yes This made me wonder about a tangential statistic that would, in all likelihood, be impossible to derive: If we looked at all database systems running at any given time, what proportion does each technology represent (e.g., Postgres vs. MySQL vs. [your favorite DB])? You could try to measure this in a few ways: bytes written/read, total rows, dollars of revenue served, etc. It would be very challenging to land on a widely agreeable definition. We'd quickly get into the territory of what counts as a "database" and whether to include file systems, blockchains, or even paper. Still, it makes me wonder. I feel like such a question would be immensely interesting to answer. Because then we might have a better definition of "most of the time."
- abtinf 11mo agoSQLite likely dominates all other databases combined on the metrics you mentioned, I would guess by at least an order of magnitude. Server side. Client side. iOS, iPad, Mac apps. Uses in every field. Uses in aerospace. Just think for a moment that literally every photo and video taken on every iPhone (and I would assume android as well) ends up stored (either directly or sizable amounts of metadata) in a SQLite db.
- sublimefire 11mo agoYes it seems like it is absent in this discussion but maybe it should have been “it” the whole time as a default option. I wonder if it could attain similar throughput numbers; bet the article would feel slightly sarcastic then though
- aussieguy1234 11mo agoI've found Kafka to be not particularly great with languages other than Java, if Confluent schemaregisty is involved. I had fun working with the schema registy from TypeScript.
- udave 11mo agoI find the distinction between queue and pub sub system quite poor. A pub sub system is just a persistent queue at its core, the only distinction is you have multiple queues for each subscriber, hence multiple readers. everything else stays the same. Ordering is expected to be strict in both cases. The Durability factor is also baked in both systems. On the question of bounded and unbounded queue: does not message queues also spill to disk in order to prevent OOM scenarios?
- woile 11mo agoThere are a few things missing I think. I think kafka makes easy to create an event driven architecture. This is particularly useful when you have many teams. They are properly isolated from each other. And with many teams, another problem comes, there's no guarantee that queries are gonna be properly written, then postgres' performance may be hindered. Given this, I think using Kafka in companies with many teams can be useful, even if the data they move is not insanely big.
- lmm 11mo agoIf Kakfa had come first, no-one would ever pick Postgres. Yes, it offers a lot of fancy functionality. But most of that functionality is overengineered stuff you don't need, and/or causes more problems than it solves (e.g. transactions sound great until you have to deal with the deadlocks and realise they don't actually help you solve any business problems). Meanwhile with no true master-master HA in the base system you have to use a single point of failure server or a flaky (and probably expensive) third-party addon. Just use Kafka. Even if you don't need speed or scalability, it's reliable, resilient, simple and well-factored, and gives you far fewer opportunities to architect your system wrong and paint yourself into a corner than Postgres does.
- dagss 11mo agoI really believe this is the way: Event log tables in SQL. I have been doing it a lot. A downside is the lack of tooling client side. For many using Kafka is worth it simply for the tooling in libraries consumer side. If you just want to write an event handler function there is a lot of boilerplate to manage around it. (Persisting read cursors etc) We introduced a company standard for one service pulling events from another service that fit well together with events stored in SQL. https://github.com/vippsas/feedapi-spec https://github.com/vippsas/feedapi-spec Nowhere close to Kafka's maturity in client side tooling but it is an approach for how a library stack could be built on top making this convenient and have the same library toolset support many storage engines. (On the server/storage side, Postgres is of course as mature as Kafka...)
- hyperbolablabla 11mo agoI for one really dislike Kafka and this looks like a great alternative
- moring 11mo agoI'll soon get to make technology choices for a project (context: we need an MQTT broker) and Kafka is one of the options, but I have zero experience with it. Aside from the obivous red flag that is using something for the first time in a real project, what is it that you dislike about Kafka?
- NortySpock 11mo agoNote: by "client" I mean "consuming application reading from a Kafka topic" Not your parent poster, but Kafka is often treated like a message broker and it ain't that. Specifically, it has no concept of NACK-ing messages, it is either processed or not processed. There's no way to the client to say "skip this message and hand it to another worker" or "I have this weird message but I don't know how to process it, can you take it back?". What people very commonly do is to instead move the unprocessed message to a dead-letter-queue, which at least clears the upstream queue but means you have to sift through the dead-letter-queue and figure out how to rescue messages. Also people often think "I can read 100 messages in a batch and handle them individually in the client" while not considering that if some of the messages fail to send (or crash the client, losing the entire batch), Kafka isn't monitoring to say "hey you haven't verified that message 12 and 94 got processed correctly, do you want to keep working on them or should I assign them to someone else?" Basically, in Kafka, the offset pointer should only be incremented after the client is 100% sure it is done with the message and the output has been written to durable storage if you care about the outcome. Otherwise you risk "skipping" messages because the client crashed or otherwise burped when trying to process the message. Also Kafka topic partitions are semi-parallel streams that are not necessarily time ordered relative to each other... It's just another pinch point. Consider exploring NATS Jetstream and its MQTT 3.1.1 mode and see if it suits your MQTT needs? Also I love Bento for declarative robust streaming ETL.
- asah 11mo ago"500 KB/s workload should not use Kafka" - yyyy!!! indeed, I'm running 5MBps logging system through a single node RDS instance costing <$1000/mon (plus 2x for failover). There's easily 4-10x headroom for growth by paying AWS more money and 3-5x+ savings by optimizing the data structure.
- EdwardDiego 11mo agoI've always said, don't even think about Kafka until you're into MiB/s territory. It's a complex piece of software that solves a complex problem, but there's many trade-offs, so only use it when you need to.
- 0xDEAFBEAD 11mo agoWhy does it matter how many distinct tools you use? It seems easiest to just always use the most standard tool in the most standard way, to minimize the amount of custom code you have to write.
- sherinjosephroy 11mo ago[flagged]
- suyash 11mo agoPostgres isn't ideal, you need a timeseries database for streaming data.
- LinXitoW 11mo agoIsn't one gigantic advantage with Postgres the ACID part? It seems to me that the hardest part of going for a MQ/distributed log like Kafka is re-working existing code to now handle the lack of ACID stuff. Things that are trivial with Postgres, like exactly once delivery, are huge undertakings without ACID. Personally, I don't have much experience with this, so maybe I'm just missing something?
- mrkeen 11mo agoIt is a gigantic advantage! And you are missing something! ACID is an aspirational ideal - not something that 'just works' if you have a database that calls itself ACID. What ACID promises is essentially "single-threaded thinking will work in a multi-threaded environment." Here's a list of ways it falls short: 1) Settings: Postgres is 'Read Committed' by default (which is not quite full Isolation). You could change this, but you might not like the resulting performance drop, (and it's probably not up to you unless you're the company DBA or something.) 2) ACID=Single-node-only. Maybe some of the big players (Google?) have worked around this (Spanner?), but for your use case, the scope of a (correct) ACID transaction is essentially what you can stuff into a single SQL string and hand to a single database. It won't span to the frontend, or any partners you have, REST calls, etc. It's definitely useful to be able to make your single node transition all-or-nothing from valid state to valid state, but you still have all your distributed thinking to do (Two generals, CAP, exactly-once-delivery, idempotency, etc.) without help from ACID. 3) You can easily break ACID at the programming language level. Let's say you intend to add 10 to a row. If you do a SELECT, add 10 to the result, and then do an update, your transaction won't do what you intended. If the value was 3 when you read it, all the database will see is you setting the value to 13. I don't know whether the db will throw an exception, or retry writing 13, but neither of those is 'just increment by 10'. The reason I use Kafka is because it actually helps with distributed systems. We can't beat CAP, but if we want to have AP, we can at least have some 'eventual consistency', that is, your services won't be in exact lockstep from valid-state to valid-state (as a group), but if you give them the same facts, then they can at least end up in the same state. And that's what Kafka's for: you append facts onto it, then each service (which may in fact have its own ACID DB!) can move from valid-state to valid-state (even if external observers can see that one service is ahead of another one).
- coldtea 11mo ago>The claim is that it handles 80%+ of their use cases with 20% of the development effort. (Pareto Principle) The Pareto principle is not some guarantee applicable to everything and anything saying that any X will handle 80% of some other thing's use cases with 20% the effort. One can see how irrelevant its invocation is if we reverse: does Kafka also handle 80% of what Postgres does with 20% the effort? If not, what makes Postgres especially the "Pareto 80%" one in this comparison? Did Vilfredo Pareto had Postgres specifically in mind when forming the principle? Pareto principle concerns situations where power-law distributions emerge. Not arbitrary server software comparisons. Just say Postgres covers a lot of use cases people mindlessly go to shiny new software for that they don't really need, and is more battled tested, mature, and widely supported. The Pareto principle is a red herring.
- MrDarcy 11mo agoIs the mapping of use cases to software functionality not a power law distribution? Meaning there are a few use cases that have a disproportionate affect on the desired outcome if provided by the software?
- ses1984 11mo agoYou might be right, but does anyone have data to support that hypothesis?
- ploxiln 11mo agoIt probably applies better to users of software, e.g. 80% of users use just 20% of the features in Postgres (or MS Word). This probably only works, roughly, when the number of features is very large and the number of users is very large, and it's still very very rough, kinda obviously. (It could well be 80% / 5% in these cases!) For very simple software, most users use all the features. For very specialized software, there's very few users, and they use all the features. > The claim is that it handles 80%+ of their use cases with 20% of the development effort. (Pareto Principle) This is different units entirely! Development effort? How is this the Pareto Principle at all? (To the GP's point, would "ls" cover 80% of the use cases of "cut" with 20% of the effort? Or would MS Word cover 80% of the use cases of postgresql with 20% of the effort? Because the scientific Pareto Principle tells us so?) Hey, it's really not important, just an idea that with Postgres you can cover a lot of use cases with a lot less effort than configuring/maintaining a Kafka cluster on the side, and that's plausible. It's just that some "nerds" who care about being "technically correct" object to using the term "pareto principle" to sound scientific here, that bit is just nonsense.
- dev_l1x_be 11mo agoApples are sweet, I am going to eat an onion. I love these articles. > The other camp chases common sense It is never too late to inject some tribalism into any discussion. > Trend 1 - the “Small Data” movement. 404 Just perfect.
- Nifty3929 11mo agoI do agree that too often folks are looking for the cool new widget and looking to apply it to every problem, with fancy new "modernized" architectures and such. And Postgres is great for so much. But I think an important point to those in camp 2 (the good guys in TFA's narrative) is to use tools for problems they were designed to solve. Postgres was not designed to be a pub-sub tool. Kafka was. Don't try to build your own pub-sub solution on top of Postgres, just use one of the products that was built for that job. Another distressing trend I see is for every product to try to be everything to everyone. I do not need that. I just need your product to do it's one thing very well, and then I will use a different product for a different thing I need.
- ARandomerDude 11mo agoI'm solidly in camp 2, the "common sense" camp that doesn't care about buzzwords. That said, I don't consider running Kafka to be a headache. I work at a mid-sized company, processing billions of Kafka events per day and it's never been a problem, even locally when I'm processing hundreds of events per day. You set it up, forget about it, and it scales endlessly. You don't have to rewrite anything and it provides a nice separation layer between your system components. When starting out, you can easily run Kafka, DB, API on the same machine.
- enether 11mo agoI also strongly believe it's not a headache. Vendors frequently push that narrative so they can sell their own managed (or proprietary) solution on it. With a decent AI model (e.g ChatGPT Pro), it's easier than ever to figure out best practices and conventions. That being said, my point is more about the organizational overhead. Deploying Kafka still means you need to learn how it works, why it's good, its configs, API, how to debug it, set up obesrvability, yada yada.
- throwwgisgreat 11mo ago> processing billions of Kafka events per day Except that the burden is on all clients to coordinate to avoid processing an event more than once since Kakfa is a brainless invention just dumping data forever into a serial log.
- williamdclt 11mo agoI'm not sure what you're talking about. Do you mean different consumers within the same consumer group? There's no technology out there that will guarantee exactly-once delivery, it's simply impossible in a world where networks aren't magically 100% reliable. SQS, RedPanda, RabbitMQ, NATS... you call it, your client will always need idempotency.
- mrkeen 11mo agoThat is called a 'consumer group' which has been a part of Kafka for 15 years. The author is suggesting to avoid this solution and roll your own instead.
- GrumpyGoblin 11mo agoThere is another aspect that many people aren't discussing, the communication aspect. For a medium to large organization with independent programs that need to talk to each other, Kafka provides an essential capability that would be much slower and higher risk with Postgres. Standardizing the flow of information across an organization is difficult. Kafka is crucial for that. To achieve that in Postgres would require either a shared database which is inherently risky or would require a customized API for access which introduces another layer of performance bottleneck and build/maintenance cost and decreases development productivity/performance. So you have a double whammy of performance degradation with an API. And for multiple consumers operating against the same events (for example: write to storage, perform action, send to data lake), with a database you need a magnitude more access, so N*X with N being the number of consumers multiplied by the query to consume. With three consumers you're tripling your database queries, which adds up fast across topics. Now you need to start fixing indexes and creating views and other workload to keep performance optimal. And at some point you're just poorly recreating Kafka in a database. The common denominator in every "which is better" debate is always use case. This article seems like it would primariy apply to small organizations or limited consumer need. And yea, at that point why are you using events in the first place? Use a single API or database and be done with it. This is where the buzzword thing is relevant. If you're using Kafka for your single team, single database, small organization, it's overkill. Side note: Someone mentioned Postgres as an audit log. Oh god. Done it. It was a nightmare. Ended up migrating to pub/sub with long-term storage in Mongo. which solved significant performance issues. Audit log is inheritently write once read many. There is no advantage to storing in a relational database.
- natmaka 11mo agoIMHO the main difference between PostgreSQL and any 'competitor' is that in most cases a software developer will quickly find not only how to use it quite properly for his use case but also why some way he adopted isn't right and triggers some non-negligible problem. There are many reasons for this: most software developers have more than a vague idea about its underlying concepts, most error messages are clear, the documentation is superb, there are many ways to tap into the vast knowledge of a huge and growing community...
- BinaryIgor 11mo agoThat's golden: "2. The other camp chases common sense This camp is far more pragmatic. They strip away unnecessary complexity and steer clear of overengineered solutions. They reason from first principles before making technology choices. They resist marketing hype and approach vendor claims with healthy skepticism." We should definitely apply Occam's razor as the industry far more often; simple tech stacks are better to manage and especially master (which you must do, once it's no longer a toy app). Introduce a new component into your system only if it provides functionality you cannot get with reasonable effort, using what you already have.
- redbell 11mo agoSee also: Redis is fast – I'll cache in Postgres: https://news.ycombinator.com/item?id=45380699 https://news.ycombinator.com/item?id=45380699
- charles_f 11mo agoNo-one ever regretted using a piece of infra to do what it wasn't planned for, during an incident at 3AM
- teleforce 11mo ago>There is also a general expectation that there is strict order - events should be read in the same order that they arrived in the system. This is the Achilles' heel of Kafka (and Pulsar) where for all streaming systems (or workarounds) with per message key acknowledgements incur O(n^2) costs in either computation, bandwidth, or storage per n messages [1]. [1] What If We Could Rebuild Kafka from Scratch?: (Top comment of the 220 comments) https://news.ycombinator.com/item?id=43790420 https://news.ycombinator.com/item?id=43790420