5 ms·
8000 messages a day, tops? That’s 5 a minute. Does that warrant “infrastructure”? I think a gameboy’s Z80 could handle that load. I don’t want to be dismissive
by SanderNL 3y ago
8000 messages a day, tops? That’s 5 a minute. Does that warrant “infrastructure”? I think a gameboy’s Z80 could handle that load.
I don’t want to be dismissive, but I often see these big numbers being posted, like “14M messages” or “thousands of messages” and then adding “per year” or something, which brings it down to toy level load.
Even the first “serious” example is about “thousands of messages” per minute. Say 5K a minute. That’s 83 per second, say 100. That seems .. not that interesting?
Am I being too dismissive? I think I am. I am not seeing something right. Can anybody say something to widen my perspective?
- wayfinder 3y agoYeah that’s nothing. My busy discussion forum built in PHP running on a toaster of a server was handling way more load than that 15 years ago.
- baz00 3y agoI think you're probably being slightly dismissive. It's not necessarily about load but various other concerns like durability, delivery latency and how failures are handled. There's a big difference between messaging and reliable messaging. I have messaging systems that take 10-20 messages a day but must deliver those messages and do it on a deadline. For that you do need infrastructure (and no that isn't a queue inside a SQL database).
- slt2021 3y agoyou don't need "systems" for 10-20 messages a day. it all could be replaced with S3 buckets and aws-cli with even better durability and delivery latency and error handling than anything you would be able to engineer yourself
- anthonyskipper 3y agoYour point is good, but that stack wouldn't win any latency awards. Many of the people I know using kafka need latencies in the millisecond range.
- slt2021 3y agokafka is not for latency, it is used for high throughput. By design kafka shines in high throughput workload due to consumer and producer concurrency (consumer group), broker concurrency (multiple nodes and partitions). for latency sensitive you will probably need redis pub/sub or something in-memory
- KaiserPro 3y agobut kafka isn't fast. Most things are backed by real files, so when you hit limits or something ejected from cache, it gets slow real fast. Kafka isn't the right choice for most things. SQS, MQTT, NATS, rabbit if you're wanting a lot of admin are all better (plus the crap that azure and google make)
- happymellon 3y agoIn all the times I've been forced to use Kafka, I have never seen single digit millisecond latency. If you need fast response Kafka is a bad choice. If you are okay going to multiple digits of milliseconds then there are simpler solutions. The only reason to use Kafka is the ability to guarantee order. For everything else it's second place at best.
- pharmakom 3y agowhat if you want to have at least once processing and durability?
- happymellon 3y agoLike AWS SQS? Others provide at least once processing as well. Kafka has been the slowest out of them that I've used and definitely more complex to use.
- pharmakom 3y agoSQS has durability for up to 14 days and you can only have one consumer group. It’s also proprietary.
- The_Colonel 3y agoThen you have dependency on a specific proprietary API / technology available from a single company. Doesn't look like a good trade off.
- slt2021 3y agoS3 API has become lingua franca and is supported by open source (minio), storage vendors (QNAP), as well as plugins that translate S3 API calls to APIs for competing cloud providers (s3proxy). all this is done because S3 provides unmatched durability and reliability at a dirt cheap cost of $22/terabyte/month of storage (with the first 50Tb/mo free!). Try to beat that reliability guarantees with whatever you handrolled, and I bet you will never be able to beat the cost of S3, even match the durability, reliability, availability guarantees at any reasonable cost at all from https://aws.amazon.com/s3/storage-classes/ https://aws.amazon.com/s3/storage-classes/: Key Features: Low latency and high throughput performance Designed for durability of 99.999999999% of objects across multiple Availability Zones have you ever built anything with 11 nines? (as in eleven nines)
- herbstein 3y ago> S3 API has become lingua franca S3 API support sounds great until your costumer builds a system with an "S3 compatible object storage" product. Soon you discover that many "S3 compatible" solutions aren't actually that compatible when pushed.
- slt2021 3y agoget-object and put-object is really all you need. everything else is nice to have
- baz00 3y agoS3 is fine until you want your data to leave AWS. Then it costs $92 / TB to get it out again. Also S3 has durability guarantees but it's very difficult to do a durable transactional write to S3. Try it a few million times and see. The API is a defacto shitty standard. These two facts are rather interesting when it comes to doing a restore from your supposed backup or wonder why consistency guarantees between external metadata services (DB) and what is in S3 don't always line up.
- baz00 3y agoI am completely dumbfounded by this reply. You're suggesting we engineer something on S3 and aws-cli, while complaining about engineering something ourselves when AWS offers a perfectly good queue service that requires no engineering? Uff. I'm going to buy a hut in the woods and live in it.
- slt2021 3y agoI used s3 just as an example of a service with very good track record of availability for a very low cost - and perfectly reliable available service can be created with bash scripts, aws-cli and free-tier AWS account. perfectly fine with using SQS, just it will have worse availability guarantees than S3 - people should understand tradeoffs
- Brian_K_White 3y agoI don't know but I think they might be suggesting that the answer to "don't need a huge complex system" is not "use someone else's huge complex system".
- slt2021 3y agomy understanding is that you should not roll your own expensive complex huge system, just use AWS S3 which provides eleven nines of availability at $22/tb/mo. It is really hard to beat the cost/benefit ratio of S3. a lot of mediocre engineers cannot swallow a pill that all their expensive work with hundreds of hours of overengineering could be replaced by a couple of AWS managed serverless services stitched together with a few mouse clicks or a single .yml file
- SanderNL 3y agoWhy does AWS even enter the conversation?
- catchnear4321 3y agosame reason any vendor enters the conversation. the argument that you can get more value out of the vendor than you would have out of a full time employee, and often for less money. sometimes it really does make sense just to pay someone else to solve the problem. not always, but not never.
- rewmie 3y ago> 8000 messages a day, tops? That’s 5 a minute. Does that warrant “infrastructure”? I think a gameboy’s Z80 could handle that load. The blog post is quite clear in stating that their pain points had nothing to do with scaling or throughput. The author explicitly mentions idempotency, custom headers, and authentication. I think you're ranting about a strawman you put up.
- SanderNL 3y agoI think you might be right. I’ll take the hit. Still, I think talking about Kafka without load feels like Christmas without tree. I thought that was the main point of it. Handling massive loads (by distributing them).
- rewmie 3y agoI think that the author just leveraged the existing messaging infrastructure they were already using for other services, and also reused their know-how. Their new microservice will barely register in the overall traffic volume, but they don't need to either deploy dedicated infrastructure or onboard onto yet another technology just to have message listeners.
- jshen 3y agoThe author's talked about their pain points, but they didn't really outline the downside of going to this approach in a comprehensive way. What I've found with event based systems is that they are much worse if your operational maturity/excellence is low. To make it worse, operational excellence can get worse over the years so you can't simply base the decision on how well you do it today.
- karmakaze 3y agoIt seems to me that Kafka was the strawman that the article put up to take down as an antipattern. There was no part of the premise that made sense. > While my client was in a situation where we had no choice but to use this monstrosity of Kafka, Avro, and custom message headers, I would never recommend this usage of Kafka if I had the option. The antipattern isn't Kafka, it's the client situation.
- Zandikar 3y agoFTA's conclusion: "If you are handling thousands of messages a day, a simple database-driven queue might be better than Kafka." They're not trying to say "thousands of messages a day" is a lot, but rather not. Or at the very least, they're saying that at that scale, it is not significant enough to merit the complexity they were dealing with.