12 ms·
What I wish someone would have told me about using RabbitMQ (2020)
- gigatexal 4y agoWell written and harrowing tale of the CAP theorem striking at 4:45AM.
- Jedd 4y agoNeeds a [2020], but that suggests the author may well have an answer by now to their hypothetical question 'how do you propose to upgrade the cluster?'. Like a lot of highly complex Cool Tools they're marvellous right up until you hit an edge case or performance threshold or an odd failure state -- and then you find yourself copy-pasting increasingly trimmed-down log entries, desperately seeking people who've hit the same problem, or rather, people who've solved the same problem and thought to describe it on the Internet. If you're on fresh software, a fresh version, or just doing something mildly off-label, this can be a despairing process.
- deleted 4y ago[deleted]
- beckingz 4y agoClassic issues with distributed systems.
- isoprophlex 4y agoWell written story... also: > I have personally experienced network partions happening in two ways: all nodes in the cluster being updated at the same time through Windows update and firewall rules. I'd rather try to build a four story brick building with horse dung for masonry mortar, than run my backend on windows boxes. How many cumulative hours of horror has Windows Update unleashed on the human race?
- logifail 4y ago> I'd rather try to build a four story brick building with horse dung for masonry mortar, than run my backend on windows boxes. How many cumulative hours of horror has Windows Update unleashed on the human race? +1 War story: an ex-client of mine was on a different continent demoing a product (involving a physical device talking to a software backend hosted by a partner company) at a neutral site to a potential new customer. Months and months of talks and prep, dry run demos in the weeks before, everyone was quietly confident. Scope for selling lots and lots of these devices if the demo goes well, but the marketspace is fairly ruthless. Minutes before the team are due to present, a sudden and total comms loss between the device and the backend. Cue frantic calls which due to the time zones even involved getting the CEO of the partner company out of bed. Minutes ticking by.. Potential customer not impressed, said: you have ten minutes, after that we're leaving. As the potential customer was walking out of the building, the partner company finally tracked down the issue. Windows Update, on a backend database server.
- buran77 4y agoIn all honesty you should rarely (if ever) patch or reboot all Prod servers in a single wave. No matter how much testing you think you've done in your development or pre-production environments something will still surprise you in prod and that costs more. Always and without exception, no matter what OS or software you use, make sure that after a patching cycle (or impactful change) you have a chunk of your environment untouched and still fully functioning. There will be situations where you can't, special cases where you're left with few options. If the system is that critical have 2 of everything (clusters, databases, etc.) in a way where the mirror can stay untouched and keep your business going. And if it needs saying, something like updates or reboots should never happen truly automatically (as in the software decided when or if to do it). A human should decide how to automate that and have full control over when and if it happens.
- justsomehnguy 4y agoI don't remember any server systems to reboot for updates if not configured to do so. And I've seen enough WinSvr in my life. Probably nobody consulted a sysadmin when they designed the system. *grin* Windows 10 behaviour is another story, but I would always stand on what it's a self-inflicted wound - people blamed MS for not updating soon enough... while disabling or postponing updates for literally years. MS acted. Poorly, yes, because by that time it ascended to General Motors levels of bureaucracy[0], but again - this is another story. [0] https://www.joelonsoftware.com/2006/06/16/my-first-billg-review/ https://www.joelonsoftware.com/2006/06/16/my-first-billg-rev...
- quantum_mcts 4y agoStopped reading when saw "Windows" to check the comments. Here it is.
- hinkley 4y agoI was part of a pretty large meetup for many years, and all through the dotcom boom and the following years the Windows Laptop behavior was utter trash. Any time a presenter showed up we had to fuck around for ten minutes to get their projector settings working, and for some reason I was the only one who could remember to ask them to turn off their screen saver. If I came late or someone else volunteered to help, we’d get halfway through their presentation, they would launch into a long anecdote or answer a question from the audience, the screen saver or sleep mode would kick in, and when they unlocked they would discover that Windows had completely forget the projector settings, and we’d have to stop for five minutes in the middle of a thought in order to fix it. There are so many other, better reasons to say it besides this or automatic updates, but Fuck Microsoft. Right in the ear.
- radicalbyte 4y agoIt's configurable and particularly if you're running a production system you'll be wanting your own domain controller configured correctly to control Windows Update. If you're using the OS version which is designed to by used on Grandma's desktop as a server then you should expect a world of pain. Me? I try to avoid using Windows hosts because .Net runs great under Linux and you don't need to have 5 different Windows VMs running on your dev machine to have an accurate development environment.
- jiggawatts 4y agoWhat does this have to do with Windows? Mass-reboots or network outages happen for all sorts of reasons. Power outages happen. Network switches fail. Firewalls get misconfigured. Scripts get mis-scheduled. Windows Update is just one of many ways I've seen all-at-once reboots of servers. I don't see how this would be better on Linux. Do Linux admins never make mistakes in cron jobs? Is Linux magically immune to switch hardware failures? If RabbitMQ couldn't tolerate reboots, then that's on RabbitMQ, not Windows. PS: SQL Server AlwaysOn Availability Groups will happily restart after an all-servers simultaneous reboot. ZooKeeper won't. By "design", apparently. It converts a temporary outage into a permanent one... on purpose.
- chrisandchris 4y agoThe major difference between the examples you mention and Windows Update is that your examples are _failures_ while Windows Update is considered a feature. I usually don't "apply router failures" while I do apply updates to Windows. When was the last time a switch just restarted to apply updates? i do not dislike Windows Update but I want to control it and apply things that are convenient to me.
- jiggawatts 4y ago> When was the last time a switch just restarted to apply updates? When it's updated by a network team. It's a semi-regular occurrence in many workplaces. Windows Update is not forced upon Server editions. It can be scheduled at will by administrators. This complaint is the equivalent of saying that Linux is terrible because an admin scheduled "sudo apt-get update" to run on every node of a cluster at the same instant. Why would admins running updates concurrently be a failure of Windows, but not a failure of Linux? I seriously don't get this complaint, ESPECIALLY because the article makes it clear that: 1. The fault was in RabbitMQ, because I quote: "The default partition handling strategy is ignore which means to just enter the partitioned state and keep trucking along in this “split brain” mode thereby thrusting your cluster in to total chaos." [1] 2. Windows update wasn't inherently at fault: "The fix for Windows update was the ensure that nodes in the cluster are patched at different times." I bet the default partition configuration setting of RabbitMQ is also wrong when running on Linux, and the behaviour would be identically bad if an admin scheduled updates to run concurrently on all cluster nodes. [1] That's clown-shoes programming, which is why I'm never using RabbitMQ for anything ever. I've only heard bad things, like "the default settings result in corruption and data loss."
- what-the-grump 4y agoThere is a thing called high availability, we do it in Linux, we do it in windows, we do it in microservices, PaaS deployments. Patching and rebooting all your nodes in a cluster at once is not highly available. But hey, let’s focus on windows hate instead.
- cies 4y ago> But hey, let’s focus on windows hate instead. I think GP acutely points out that one of the problems the author has with RabbitMQ is actually a problem with (a misconfiguration of) their underlying OS, MS Windows. If you've been bitten by this, I understand the hate completely.
- yonixw 4y agoYou are correct, I woke up one day to 5 different servers in different env not responding. Across clouds (AWS+Azure). Total meltdown, are we hacked? Is the cloud down? Nope. Just ubuntu deploying docker update that had problem restarting the docker service again. Simple restart fixed all of them. But I guess people do have some bias, since on PC, windows updates are really predatory.
- smileybarry 4y agoIf you use Windows Server and not just repurpose a client SKU, you get full control over which updates, schedule, distribution, restart times, etc. IIRC it even ships in "update manually" mode, you have to manually enable automatic updates. And if you set-up something like WSUS I think you can even set gradual update rollout with health checks.
- iasay 4y agoSort of. Sometimes it goes wrong. I've seen a primary SQL Server go down because WUA took up enough CPU and interrupts to stop the cluster service health pings leading to a BSOD. This happened outside any schedule.
- smileybarry 4y agoGuessing this was in the XP-2003 era if you mentioned WUA? Windows Update had a couple rewrites in Windows 7 & Windows 8.1 to solve the old "checking for updates takes too much CPU % if you have X+ updates installed". It's dramatically better now.
- iasay 4y agoThis was last year on windows server 2019. I say WUA. I don’t know what the current version of it is called. There are still problems.
- smileybarry 4y agoAh, I see. I just remember WUA from that one EXE people ran on Windows XP to manually check for updates.
- KingOfCoders 4y agoI need to add this story: Our Oracle database on Windows NT was fast then always slow (in the 90s, we could not afford an IBM/SUN for Oracle). When I went to the computer it was fast again. Back at my desk it got slow. Reason: Pipes software rendered OpenGL screen saver.
- leephillips 4y ago“Well written”—with a horrible grammatical error right in the title.
- benjaminwootton 4y agoMessaging platforms have been an ever present in my career. Tibco, IBM MQ, back to Tibco, RabbitMQ then Kafka. I like how on the surface they can be quite simple to use, but the optimisation and management of them can be fiendishly subtle. Some of my most interesting projects have been trying to squeeze more messages through a pipe with lower latency, changing the way messages are sent and flow through these platforms, or digging into why one in a billion messages are dropped. There was also an interesting phase of trying to containerise and infra-as-code Kafka. It’s all like plumbing for data infrastructure. An interesting corner of the IT industry.
- hardwaresofton 4y agoDon't forget the NATS-based solutions - NATS Core[0] as an ephemeral message exchange (personally what I would use RabbitMQ for) - NATS Jeststream[1] as a persistent, queue-focused kafka alternative - Liftbridge as an alternate implementation of a persistent kafka alternative[2] Liftbridge has a decent comparison page[3] but unfortunately it's still missing NATS JetStream[4]. I want to see more members of the community use and write about the NATS ecosystem -- I rarely hear complaints. [0]: https://docs.nats.io/nats-concepts/core-nats https://docs.nats.io/nats-concepts/core-nats [1]: https://docs.nats.io/nats-concepts/jetstream https://docs.nats.io/nats-concepts/jetstream [2]: https://liftbridge.io/ https://liftbridge.io/ [3]: https://liftbridge.io/docs/feature-comparison.html https://liftbridge.io/docs/feature-comparison.html [4]: https://github.com/liftbridge-io/liftbridge/issues/104 https://github.com/liftbridge-io/liftbridge/issues/104
- TedDoesntTalk 4y agoThere are dozens of other solutions you are leaving out. There was an entire industry around “enterprise messaging bus and orchestration“ (or a similar name) in Java land in the early 2000s.
- hardwaresofton 4y agoRight, and with all due respect I'd like to leave those in the past which is why I didn't mention not forgetting them! NATS is a lively project, has never had anything to do with mules or camels, and isn't hard-tied to concepts like ESB. It deserves to be mentioned in 2022. But here are some projects I left out that do deserve to be mentioned I think: - RedPanda[0] - Apache Pulsar[1] [0]: https://redpanda.com/ https://redpanda.com/ [1]: https://pulsar.apache.org/ https://pulsar.apache.org/
- kelnos 4y agoI get that distributed messaging/queuing is difficult (been there, done that, often didn't do a great job of it), but the constraint that every node in the cluster has to be running the exact same version of RabbitMQ is ridiculous. I can't see how you could ever orchestrate zero-downtime upgrades. Requiring that they're all on the same major version sounds reasonable, with a further constraint that clients can only use features supported by every node in the cluster (e.g. if a new feature was introduced in 1.3.0, and some nodes are still running 1.2.x, clients shouldn't use that new feature until all nodes have been upgraded). And there should still be some sort of reasonable migration process to the next major version! It may not be a simple migration process, but should at least be something where you can orchestrate things such that you have no downtime. Having the default behavior during a network partition be "whatever, just chug along as if nothing is wrong" is bonkers. Yes, the person who first set up the cluster should have read the documentation and gone over the configuration file line by line to see what might need changing, but... damn, that's a terrible default. Sure, some people's applications might value availability over consistency, but that's not the safest choice that follows the principle of least surprise. Using a higher-level library to interact with the cluster is really good advice, in general. We used Kafka at my last company, and colleagues who actually knew what they were doing wrote a (simple) wrapper library that set things up properly for our cluster so clueless people (such as myself) could write producers and consumers without having to understand what all the fiddly connection setup settings did, and how to handle various edge-case errors. Before that, quite a few outages were due to producer/consumer misconfigurations. Also I can't imagine running this kind of infra on Windows servers. That sounds like a self-inflicted wound (by "self" I mean the company, not OP specifically). And the idea of Windows Update running on a prod server ruining your day... what? IMO infra should be as immutable as possible. Patching/updating software on a machine should be a matter of spinning up a new machine with an already-updated image (that you've built and tested elsewhere), bringing those new machines into a cluster, and then decommissioning the old ones. When colos and dedicated servers were all the rage, that was difficult (and sometimes impractical), but in this day and age, with on-demand cloud provisioning, there's no excuse for companies that can use that sort of infra.
- vorpalhex 4y agoYou do zero downtime upgrades via cluster federation. Works great.
- jazdw 4y agoIs publishing a request to a queue and then polling for a response a typical pattern for a distributed web application?
- jhh 4y agoI was also surprised by that but there may be some slightly unusual background reason for that that simply doesn't get mentioned. Superficially it sounds like the author should just... do the HTTP request? Since they're waiting for the response anyway
- physicsguy 4y agoI'm not sure the exact idea here because they said something about retrieving PDF and JSON data, but a big reason for not doing this is normally that some calculation or whatever takes a lot longer than you want a HTTP request to take. So you effectively push the work off, and then the user comes back at a later time and something has been processed and they can see the result.
- poulsbohemian 4y agoAlmost twenty years ago, I was on a team that built a bunch of health care apps where we had significant, robust mainframe business systems that we needed to get in front of both internal staff, external providers, and the public. So yes, those systems definitely used a whole lot of messaging and in my career since I saw repeats of those patterns not only in that industry, but also financial services, telco, and lots of other places I stopped along the way. Anyplace you’ve got hulking business systems with lots of rules, it’s common to have messaging in between there and the web, and for various distributed applications to talk to each other.
- nightpool 4y agoNit: the application the author describe doesn't even publish the request using RabbitMQ, as far as I can tell—they publish the request with a simple HTTP request. They use RabbitMQ solely for scheduling the "Check if this task is done yet" job every ~5 minutes. I haven't used RabbitMQ myself, but it does sound like kind of like a square peg / round hole situation to me.
- jillesvangurp 4y agoI'd say this is a good introduction into reasons why self hosting clustered software is not something to do lightly. He's basically running rabbit mq on Windows. Yikes. And also, why? I'm sure it can be done responsibly. But allowing all nodes to self update and reboot randomly sounds like amateur hour to me (that actually happened apparently). So, don't do that. I don't use rabbit mq currently but we do use Elasticsearch, which has similar clustering capability and used to be more susceptible to split brain situations (been there, dealt with that) These days, I recommend using Elastic Cloud and avoid self hosting it. It's only cheaper until the first time you have to deal with a split brain cluster because you botched an update, mis-configured it, etc. One of the nice features in Elastic cloud is that you can click an update button and it will orchestrate a rolling restart. If you don't know what that is, you should not be operating a cluster of any kind. I'm sure there are similar cloud based services for rabbitmq. Probably well worth the money unless your in house ops team is super experienced with operating it (which clearly wasn't the case here). Such a team would cost you many hundreds of thousands of dollars per year. One person does not cut it. You need at least a 3 or 4 so you can afford them taking vacations, sick leave, leaving, or dying in some tragic accident. Half a million pays for some pretty nice cloud based clustering capacity. A good team will cost you more. For reference, we pay about 70/month for a tiny Elastic cloud search cluster that is actually good enough. I can double the price and capacity with a simple slider and it would still be cheap. One hour of my time is more than that with my normal freelance rate. My monthly rate would pay for an enormous cluster that far exceeds anything we need and it would still be cheaper than making sure we have four people with my skill set in the team (we don't) able and willing to look after it at all hours. The largest cluster I've ever dealt with was a self hosted monster that could index billions of documents per hour (millions per second). You so much as looked wrong at it, all hell would break loose. That cluster was one of several managed by a very experienced ops team that probably cost millions per years. That's the price of doing business at scale when self hosting. Most companies running into trouble with ES cut corners on doing it right and then pay the price with preventable outages, scaling issues, technical debt, etc. Cloud based services don't completely prevent this but if you know what you are doing, they provide a nice level of safety and risk mitigation.
- vladvasiliu 4y ago
- brundolf 4y agoGenuine question from someone who doesn't know any better: what's the advantage of having a message queue like this (service publishes message to queue, recipient gets notified about it, responds to payload) vs just sending an HTTP request directly to the recipient?
- avmich 4y agoQueues are used to be able to have bursts of requests without overloading the server, which may be less performant than needed to service those bursts of requests directly. With the message bus the clients can (roughly) send requests as frequently as they want, while servers will handle them as fast as they can, without the danger of missing a request or having clients to wait. The queue is the "asynchronization" mechanism.
- deleted 4y ago[deleted]
- vladvasiliu 4y agoNot an expert in this domain, but off the top of my head: It's easier to handle certain situations, like if you have multiple recipients, you may want to notify all of them, at least one, at most one, etc. It's also a way of handling consumers that may be unavailable for whatever reason and your source may not want to have to deal with that. For example, if the producer is a cron job, you may not want it to have to hang around for an arbitrary time if your consumers are already busy. Sure, this means that the queue is available, but in principle it's supposed to be more available than the consumers.
- 66fm472tjy7 4y agoWe use RMQ for most of our asynchronous processing. In most cases, we get a HTTP call and publish a message to the RMQ after committing the DB transaction, then we send the response to the HTTP client. We found out the hard way that RMQ does not behave like a transactional DB. Just because publishing worked does not mean the message will be delivered. Our solution is to also write the message into an outbox table in the DB. We then publish the message using confirms[0]. RMQ asynchronously sends us a confirmation when it has really persisted the message. We then delete the outbox entry. If we do not receive the confirmation in time, a timer will re-publish the message. Therefore I disagree with the suggestion of using a library wrapping the native RMQ one. We are using spring-amqp and this made it harder to understand what is going on. In the end, for a large project you will have to understand nuances of RMQ (and other infrastructure you are using). Using a leaky abstraction over it means you now have to understand both the underlying product and the abstraction. [0] https://www.rabbitmq.com/confirms.html#publisher-confirms https://www.rabbitmq.com/confirms.html#publisher-confirms
- deleted 4y ago[deleted]
- cmckn 4y agoHow quickly does RMQ ack the message? Obviously too long to delay an HTTP response, or you’d have skipped the DB part of this; but this seems kind of clunky. I know Kafka has (optional, tunable) acknowledgements for publication, for example, that you could use for this.
- 66fm472tjy7 4y agoIn the first iteration of using confirms, we did not have the outbox but only logged how long it took to get the confirmation. After 3 seconds, we would throw out the expected confirmation. If a confirmation took longer than that, we would log that we received an unknown confirmation. We hoped it would be fast enough that we can just wait for the confirmation before committing the transaction. The official documentation says > This means that under a constant load, latency for basic.ack can reach a few hundred milliseconds I never did statistics, just looked at the log. IIRC most were acceptable but > 3s occurred frequently enough (and we even had instances of messages never being confirmed, IIRC) that we abandoned that plan. We considered using Debezium[0], but decided on the current solution as it could be solved entirely with the current services and infrastructure whereas Debezium would have required us to deploy (writing this from memory so this might be inaccurate/incomplete) Kafka, Zookeeper, and a connector service. [0] https://debezium.io/ https://debezium.io/
- doommius 4y agoYup. Did my masters thesis on Jepsen tests and there's a huge discrepancy on how most people perceive databases, consistency models and their actual behavior. Also just driver implementations, cluster configuration and other oddities that might bring chaos and disaster out of nowhere.
- vimwizard 4y agoRedpanda talk was pretty interesting. Apparently some documentation on certain client settings is flat out wrong regarding consistency guarantees, IIRC written by Confluent.
- northstar702 4y agois your thesis available to read? Jepsen and these concepts for databases are certainly very different from those for kafa, as we heard from Kyle during the Jepsen testing for Redpanda. Someone needs to write about those perceptions!
- codedokode 4y agoHere is what I dislike about RabbitMQ: by default it opens several ports on external interfaces and doesn't provide an easy way to bind them to localhost only. To be specific, those are epmd port 4369, and RabbitMQ's ports 5672, 15672, and 25672. With help of Google I managed to move three of those ports to localhost, however the cluster communications port 25672 cannot be bound to localhost only. I found an issue on Github and developers' point is that it doesn't make sense to build a cluster from a single node, so you must have this port exposed to whole Internet so that any random stranger can connect to your RabbitMQ instance [1]. So it seems that RabbitMQ wants to advertise how scalable it is by enabling clustering by default and accepting connections from anyone even though it is not secure. It is targeted at large corporations with giant clusters and doesn't care about developers who just need a single instance. Despite my guess that most developers actually do not have a volume of messages that would justify setting up a cluster. And as RabbitMQ is written in an exotic language (Erlang) I cannot even read the code. This is the problem not only with RabbitMQ, Elasticsearch also has clustering settings enabled by default, so when you try to run two independent instances, they connect to each other and start exchanging data and creating problems without you expecting it. So annoying and so difficult to disable. In earlier versions, as I remember, they would also try to scan local network and connect to any node they could find. [1] https://github.com/rabbitmq/rabbitmq-server/issues/1661 https://github.com/rabbitmq/rabbitmq-server/issues/1661
- im3w1l 4y agoCan't you just block the port in your firewall of choice?
- codedokode 4y agoIn theory I could learn Erlang and patch the source code as well, however, I would prefer a proper solution, not a hacky workaround. It doesn't make sense if you first open the port and then block it with a firewall. Also, this is a bad idea because firewall config is kept separately from RabbitMQ config, it is easy to forget that you have a RabbitMQ instance exposed to the whole Internet and accidentally unblock the port.
- codedokode 4y ago> The typical action sequence is the user submits a request via the web application and the backend handles that message by adding a message to RabbitMQ. The consumer gets the message and makes a HTTP call to another web service to actually submit the request. From there, the polling logic takes over and subsequent messages on the queue each represent a polling attempt to retrieve the results. If a job has no results, the consumer places a message back on a queue so we can delay the next polling attempt by a (customer configurable) amount of time. This looks like an overengineering to me, unless I have missed something. For example, I don't understand this part: "the consumer gets the message and makes a HTTP call to another web service" - why cannot that web service pull the message directly? I would implement it like this: when a client submits a request, it is added into an SQL database, and RabbitMQ is used only to notify the consumer. The client gets back a secure Job Token that it can use for polling to get the status of the job. The consumer reads the job from the database, executes it and updates status in the database. The client uses polling to know when the job is done. Of course, if you don't like polling, then you can use a WebSockets daemon that would notify a client when the job is done.
- Denvercoder9 4y ago> why cannot that web service pull the message directly? It could e.g. be an external service run by a third-party, or an off-the-shelf service that's hard to add integration to, etc.
- codedokode 4y agoIndeed, I didn't think about that.
- ikiris 4y agoThis reads like a rediscovery of the SRE book step by step. distributed clusters are cool, especially messaging ones, but you have to know how to manage them or you're adding eventual points of failure, not removing them.
- mkl95 4y ago> The time is 4:45 AM. I pull it together to realize it’s a call from a number I do not know - never a good sign. I answer and it is a coworker - my peer who runs our support team that is engaged in nearly all production issues for our customers Ugh. I hope OP is making a boatload of money.
- vimwizard 4y ago>windows >developer left, project dumped on him >doesn't know the solution not likely
- samsquire 4y agoI wrote some custom Chef code that handles a seamless Erlang and RabbitMQ upgrade of a cluster of RabbitMQ reliably. Sadly it's trapped at one company. There's probably a way to write some code to drain a split brain's data but I suspect nobody has the time to write this code. You could have a process where you block writers to the minority then drain messages and enqueue them on the majority.
- lmarcos 4y agoAt previous jobs while wearing the hat of "senior software engineer", management always expected from teams the now common "you build it you run it" way of working. The seniors of the team were supposed to know (or to learn) all the nuances of the parts that composed their infra: redis, rabbitmq, postgres, go, docker, k8s, github ci/cd, oauth,... Of course, senior software engineers should also know all the stuff related to software engineering (DS, algorithms, system design, solid, etc.) I ended up burnt out. Too much to learn, and I'm not really interested in most of the infra tooling we were using (I loved redis and got to know tons of it, on the other hand I never really cared about k8s or rabbitmq). Sadly, more often than not, companies out there still expect the mentality of "you build it you run it" which implies acquiring knowledge of all the little things in your infra that can go wrong in production. Damn it.
- muxator 4y agoWhile your burnout is understandable, I'd see the whole situation in a positive light: instead of treating complexity as someone else's problem, this forces the team to aim for simplicity. This kind of attitude is often in charge of more senior engineers, since they are the ones who most probably understand (or have lived through) all the possible nuances on wanting to rely on so many systems. If I'd were in your shoes I would have ended up burned out too if I were given the responsibility of having to understand it all, but no mandate to simplify it as much as possible.
- waynesonfire 4y agoOh yeah, this trend is toxic. It's a cost-saving technique. GraphQL is another toxic technology that relieves the responsibility of designing the API to users instead of the application owners. May as well just give folks access to your DB and have them run SQL queries themselves. Look! No need for a UI. All this at the expense of the software engineer that has to manage all this complexity themselves.
- ikiris 4y ago... who do you think is supposed to manage all this complexity if not the people who are building the software?
- 4y ago
- danbulant 4y agoAs for the last point: everything that logs will eventually run out of space, be sure to have some kind of automation to rotate those logs. Had the same problem with MongoDB and QuestDB as well, and both had logs that were far bigger than the actual data in the database.
- moron4hire 4y agoI can't imagine how that is ever useful
- shudza 4y agoThis is why software engineers shouldn't handle this type of work, and why DevOps came to existence.
- rmetzler 4y agoI feel like you mix two things which kind of are related, but not really. It’s also that you don’t even state if you talk about DevOps as a philosophical approach to how software should be designed, written, and run. Or if you talk about a job title. The main issue the article is talking about that you need to think about the future and day to day operations before you start running production workloads and even think about failure modes and handling of it before you can plan your system. And it really helps to simulate something like network partitioning before you write the bigger part of your application.
- shudza 4y ago
- Sin2x 4y agoWith RabbitMQ it's also important to understand the difference between mirror and quorum queues: https://www.rabbitmq.com/quorum-queues.html https://www.rabbitmq.com/quorum-queues.html
- Sin2x 4y agoSeems like they are going to remove classic mirror queues completely in the future: https://blog.rabbitmq.com/posts/2021/08/4.0-deprecation-announcements/ https://blog.rabbitmq.com/posts/2021/08/4.0-deprecation-anno...
- mxben 4y agoI can appreciate the points about not knowing enough to engage an expert early or using a good wrapper library, etc. But this point blows my mind: > There’s this Network Partition thing, it’s kind of a big deal How could someone using RabbitMQ cluster not consider how the cluster would behave during a partition? This is exactly the kind of thing that should be tested in a safe environment before running the cluster in production. Testing for network partitions is not something one wishes someone else would have told. It is an essential responsibility for anyone in a software engineering role. Not doing some basic testing to understand partition scenarios before running a cluster (any type of cluster) in production is a disaster just waiting to happen.
- abraxas 4y ago> Throughout that time we have scaled to 200+ concurrent consumers running across a dozen virtual machines while coordinating message processing (1 queue to N consumers) and processed hundreds of millions of messages in our .NET application. That's not a level of data volume that should require any kind of distributed messaging.
- iasay 4y agoI don't think there's enough information to make that assertion there.
- abraxas 4y agoLet's be aggressive and interpret this into a worst case and assume that the author is referring to one week and 700 million messages. That's 100M messages per day or 1157 messages per second. That volume could be handled by a single DB table in an ACID way eliminating the need for a separate cluster
- iasay 4y agoDatabases aren’t queues. I’ve been burned by that a million times over the years. Also you screw your OLTP capacity there.
- abraxas 4y agoAt these volumes they can very well be. And when that is no longer tenable the time comes to introduce a streaming message system like Kafka. In my experience the use case for something like Rabbit is a small crevice between these two solutions that I have little use for
- iasay 4y agoI wouldn’t use RabbitMQ. I’d use SQS.
- gspetr 4y agoWhat he says the issue is: We should have consulted an expert before finalizing our architecture. What I see the issue is: We have exactly Nobody in our organization whose full-time job is to perform deep[0] testing. [0] https://www.developsense.com/blog/2022/01/testing-deep-and-shallow/ https://www.developsense.com/blog/2022/01/testing-deep-and-s... "Risk coverage is how thoroughly we have examined the product with respect to some model of risk." There had been no modeling of any kind of risks related to networking issues with this technology, so there is no surprise that something went terribly wrong.