11 ms·
anyone wanna share their thoughts about deploying their own messaging system vs using a messaging system provided by their cloud provider?
by kevindeasis 6y ago
anyone wanna share their thoughts about deploying their own messaging system vs using a messaging system provided by their cloud provider?
- dtech 6y agoDistributed messaging is really hard to get right. It'll seem to work fine right up until you get weird bugs and unreliability during at the worst moments. I wouldn't recommend relying primarily on something vendor-specific like Amazon SQS, but there are very good out-of-the-box tools like RabbitMQ or Kafka available. Writing your own messaging system is like writing your own database, it's the wrong choice 99.9% of the time.
- skyde 6y agoWhile I agree I see a big problem that people often pick Kafka when they really needed Rabbitmq because of the hype around Kafka and not understanding the difference between a transactional message broker ( Rabbitmq and ibm mq ...) as used in banking and distributed log store (pulsar/ Kafka)
- LgWoodenBadger 6y agoNot sure that I’d want to rely too much on Rabbit’s transactional behavior based on their limitations. https://www.rabbitmq.com/semantics.html https://www.rabbitmq.com/semantics.html
- TheBlight 6y agoAvoiding vendor lock-in is the only quasi-sane reason I can think of.
- realtalk_sp 6y agoThe GCP Pub/Sub API has largely replicated all the features you'd want out of Kafka (including Consumer Groups). The primary consideration at this point is cost. There's an inflection point in size (at some very large message volume) where it makes sense to start running your own Kafka cluster and hire a dedicated person or two to manage it. Most companies will never get anywhere close. Any project just starting out should use Pub/Sub. One thing I really like is that GCP provides emulators of Pub/Sub et al for local testing. That used to be a bit of an obstacle not too long ago. In terms of lock-in, I don't see how that applies to an AMQ. The data moving through it should only be transiently persisted, up to a week or two at most in the usual case. If you want to avoid cloud lock-in, have DB backups, use Postgres/MySQL/etc, containerize your service(s), replicate data in object storage, etc. Common sense stuff, if that's something that's of concern. Personally, I've seen "vendor lock-in" weaponized as an excuse for a lot of costly NIH bullshit. It's painful to reflect back on a project that could have involved literally a tenth of the time and pain it ended up taking because of that one choice alone.
- gigatexal 6y agoIMO lock-in fears are overblown. Build stuff out quickly and prove your idea and then refactor when you have customers and revenue.
- SpicyLemonZest 6y agoRefactoring a growing system while maintaining bug-for-bug compatibility is extraordinarily hard. Most people I know who've gone through such a migration never want to do it again.
- halbritt 6y agoMost people are simply bad at migrations.
- cglace 6y agoIf most are bad at migrations then what should most do?
- halbritt 6y agoLearn how to do migrations well. There's plenty of literature on the topic.
- SpicyLemonZest 6y agoUltimately, building a maintainable system requires that you allow for people being bad. If 10 smart engineers contribute to your system, and they each assume future developers will be good in their own distinct way, the resulting system will be so complicated that none of them are good enough to maintain it.
- halbritt 6y agoPreferably, everything that can be is loosely coupled and extensible.
- nojito 6y agoMuch much cheaper.
- dajohnson89 6y agodoes that factor in development costs?
- ex3ndr 6y agodoes developing for NATS is so much harder than for cloud provided pubsub?
- dajohnson89 6y agoConfiguration & administration of the kafka cluster is more time-consuming.
- ex3ndr 6y agoYes, but nats is dead simple
- cameronbrown 6y agoDev time would far outweigh the cost of a managed equivalent.
- nojito 6y agoDisagree. Dev time is virtually identical between the two kinds of infrastructure.
- oweiler 6y agoWe use Amazon MSK and are pretty happy with it so far.
- sz4kerto 6y agoTestability. Single biggest reason for us not going with the proprietary queue. We can create, reset, throw away queues as we want during testing, thousands times a day. Even on a laptop.
- sa46 6y agoDo you have any advice for setting up a test Kafka environment? I'd love to be able setup a lightweight in-memory Kafka for unit-y tests without going through the whole Docker compose rigamarole.
- ndmeredian 6y agoDepending on complexity and specialisation of your use that's what may work for unit testing: 1) Abstract real library with tech-agnostic interface to remove as much tech details as possible and focus on behaviour. 2) Create simplest in-memory dummy implementation for that interface, basically a mock. 3) Write a test spec for that interface behaviour (does xxx on yyy, returns zzz otherwise) and run it agains both real wrapped interface & mock to be sure they're consistent. 4) Make code to load mock in local testing env 5) Use real kafka when running it on a testing pipeline Same way you could abstract various things - e.g. once I just abstracted proprietary services and locally flushed everything into local Postgres. You can emulate multiple services there - e.g. database, somewhat-usable queue, etc. And get some extra control nice for tests (e.g. make it to throw errors on your wish or add timeout to simulate net delays). That saved a lot of time for local development for me, without dropping quality.
- EdwardDiego 6y agoIf you're using Kafka Streams it provides a TopologyTestRunner.
- gunnarmorling 6y agoCheck out the Kafka support in Testcontainers: https://www.testcontainers.org/modules/kafka/ https://www.testcontainers.org/modules/kafka/. It is container-based, but with a very simple-to-use API for spinning up Kafka in tests. If you truly want something embedded, Debezium's KafkaCluster could be interesting: https://github.com/debezium/debezium/blob/master/debezium-core/src/test/java/io/debezium/kafka/KafkaCluster.java https://github.com/debezium/debezium/blob/master/debezium-co.... It's used four our own tests of Debezium, but I'm aware of several external users. It spins up AK and ZK embedded. Disclaimer: I work on Debezium
- antoncohen 6y agoIf you are on GCP I think the choice is simple, use Cloud Pub/Sub. Extremely simple, extremely reliable, extremely performant, fairly inexpensive, multi-region (global). No maintenance, no scaling, almost no tunables, it just works. Google provides a Pub/Sub emulator for local development. I don't really buy the vendor lock-in thing for Pub/Sub-like systems. The Cloud Pub/Sub usage pattern is basically the same as Kafka, you can have a library that abstracts away the differences. There are open source libraries that do that[1]. If you ever need to switch cloud providers, or want a messaging system to span cloud providers, you can switch without changing lots of code. [1] https://github.com/google/go-cloud/tree/master/pubsub https://github.com/google/go-cloud/tree/master/pubsub
- peterhunt 6y agoI don't think it's that simple unless I misunderstand how GCP Pubsub works. I don't think GCP PubSub will give you deterministic delivery order within a partition the way Kafka will: https://cloud.google.com/pubsub/docs/ordering https://cloud.google.com/pubsub/docs/ordering
- batter 6y agoWe have Kafka and GCP pubsub. Kafka is the way to go for us. In terms of reliability, performance, load, etc.
- skyde 6y agoAre you saying GCP pubsub have bad reliability? I never used it but would be surprised if it was down for several hours each weak or would just drop your data. I would love to know your experience
- batter 6y agoOne of the teams was complaining about behavior on load. Now another team will migrate that flow to Kafka. Another team was complaining about reliability/speed. I think eventually they stopped relying on that.