8 ms·
How Kubernetes and Kafka Will Get You Fired
- metacatdud 3y agoI am happy to see people are talking about this. Everytime I am trying to point out systems are more and more complicated and do a call for simplicity I am pushed away. I heard you are not taken seriously if you don't use a well established cloud provider or similar things. Truth be told, there are not many projects you do or will work on which need this kind of things. We though cloud will help us with a lot of things but to what cost and by cost I mean stress, data protection, money etc.
- theoldlove 3y agoQuite the advertisement for AWS. I’d like more details on what the migration to native AWS services looked like
- lrobinovitch 3y agoNot OP but I wrote about migrating to MSK from EC2 here: https://theleo.zone/posts/migrating-to-msk/ https://theleo.zone/posts/migrating-to-msk/
- Proven 3y ago[dead]
- bradwood 3y agoHaving run k8s and Kafka in a previous job (I left before I got the sack) this article rings completely true. New shop: lambda and eventbridge = life is good.
- animesh 3y agolambda and eventbridge = life is good Very interesting. I am struggling to articulate this mentally. Would you please expand on this comment?
- ranguna 3y agoI'm on a stack that also uses lambdas an event bridge, I think the OP means that the stresses originating from k8s and Kafka are gone when using lambdas and event bridge.
- bradwood 3y agoIndeed, that is exactly what I mean. There are some caveats though -- eventbridge doesn't give strict ordering or exactly-once delivery, so some work is needed to make idempotent consumers, etc. But, on aggregate, it's still much much easier than managing all the infrastructure that k8s and kafka requires.
- mt42or 3y agoVery funny article, we are spending 2 people full time (on 4) trying to building on AWS services. This is really a mess and cost a lot of human resources.
- withinboredom 3y agoIt took me 4-6 weeks to get a k8s cluster up from scratch and migrate to it. I don’t do devops for a living, but I’ve been building servers and doing devops stuff for fun for nearly 20 years. So I’m not a pro, but I know what I’m doing. If your team is a bunch of software devs, you are doing the right thing… because k8s requires a bunch of knowledge you likely don’t have if you don’t have a Linux guru on the team. Even if you are using an expensive managed solution, things will go wrong and that knowledge is needed to prevent downtime.
- joshribakoff 3y agoThis doesn’t seem consistent with my experience. When you say “4-6 weeks to get a cluster up” i wonder if you actually mean to learn kubernetes and play around with deploying things, as that would make sense. I was able to install k3s in around 15 minutes and deploy my first service, myself.
- withinboredom 3y agoI mean for production, with scripts/playbooks/firewalls/rbac/storage/networking/ingress/logging/backups/etc. Yeah, you can stand up a toy cluster in matter of minutes, but that’s not the same.
- withinboredom 3y agoFor fun, I run a bare metal k8s cluster with my blog and other projects running on it. My last three nights have been fighting bugs. Bugs with volumes not attaching, nginx magically configuring itself incorrectly, and a whole bunch of other crap. This just magically started happening, but crap like this seems to happen at least once a month. It’s to the point where I spend at least one night a week babysitting the cluster. I don’t have to pay someone else to handle this, but if did, I would get rid of k8s in a heartbeat. I’ve seen a devops team of only a few people manage tens of thousands of traditional servers, but I doubt such a small team could handle a k8s cluster of the same size. I’m considering moving back to traditional architecture for my blog and other projects. K8s has been fun, but there’s too much magic everywhere.
- red-iron-pine 3y ago> I don’t have to pay someone else to handle this, but if did, I would get rid of k8s in a heartbeat. I’ve seen a devops team of only a few people manage tens of thousands of traditional servers, but I doubt such a small team could handle a k8s cluster of the same size. This has been my experience with a lot of the "we need to be cloud native! containers!" mantra in the enterprise. Some exec gets it in their head it's a good idea (and probably gets non-trivial "referral agent fees") this is a must do, and all of the young, hip developer types are happy to cheerlead it. Two years later OpEx is exploding, most of the processes haven't yet been converted to be in the cloud, and the environment isn't noticeably better or different. It sucks, just sucks in a new and more expensive way that gives you less control of your data. Seen this at 3 x F500 orgs and with multiple cloud providers, including the big 3 + one of the well known second tiers.
- thunky 3y ago> My last three nights have been fighting bugs > I spend at least one night a week babysitting the cluster ... > K8s has been fun This is why everything sucks now.
- metacatdud 3y agoHey, you can try nomad. It works nice for small-med projects. Works well with terraform and extending with nomad clients it's a breeze. I set my personal bare metals with nomad infra and never looked back.
- thrashh 3y agoI can confirm that maintaining a Kubernetes cluster is a full time job. Due to its design, there are a lot of moving parts even for the most minimum deployments. Low key I hate touching Google-created projects. On paper technically sound but in practice a guaranteed usability disaster.
- m463 3y agoI think this might be because google engineers are rewarded for inventing new things, not so much for refining or maintaining old things.
- joshribakoff 3y agoFor some people learning to change the oil on their car could be “a full time job”, but it doesn’t change the fact oil changes are a commodity i can go pay $50 for at a shop. Similarly, any cloud provider worth it’s salt provides a managed k8 cluster as a commodity these days.
- holografix 3y agoAn explanation on what was actually being so difficult about managing Kafka and K8s would have been helpful. Why didn’t the customer use EKS?
- bsaul 3y agothat also wasn't really clear in the article : did they used managed k8s and managed kafka ? in which case i don't really understand the issue, as there's nothing to manage..
- dikei 3y agoI came expecting a war story about running Kafka on K8S. Instead, I find an advertisement for AWS server-less.
- SilverBirch 3y agoI don't really think this was an advert for AWS. My guess would be that if they'd been on Azure or Google Cloud he just would've suggested that the client use the built-in products on those platforms.
- reacharavindh 3y agoNot that I want to advocate for Kubernetes, but this is just a stepping stone for another blog post in a few years "How managed services in AWS got us fired!" detailing how when AWS changed their pricing/strategies the companies had no way out of their design choices. The "vendor agnostic" approach was the right call to make at that point in time. Sure, every business is unique and some cases fit better for a hands-off (lets pay for the convenience of AWS taking care of managed offerings) approach. But, it is a fallacy to think there is no cost to pay for that decision. The cost of operating your vendor agnostic infrastructure is replaced by your team now needing to learn the intricacies of AWS such as IAM, AWS's way of networking, backups etc. Those operational needs dont just go away, they just become "easier" and more defined as AWS's way of doing it. As a consultant, one must know where to draw the line and recommend the appropriate route.
- ryan_lane 3y agoThis is easy if all you use is EC2, as K8s is effectively just a layer over EC2. Once you start using other AWS features you're locked in anyway, and tbh, what's the point of using AWS if all you're using is EC2? Choosing to be vendor agnostic is purposely choosing to need to implement every single part of a stack yourself, rather than the (usually considerably better) AWS offerings, just in case you may want to switch away in the future. It's a massive waste of money and time for nearly every business. Unless you're in the business of providing an infrastructural-level service (like heroku), you'll ship faster and cheaper by using a single cloud and going all in.
- klooney 3y ago> what's the point of using AWS if all you're using is EC2? It's very high quality, gives great audit, you can have the same control plane in a ton of regions... EC2 is fantastic.
- ryan_lane 3y agoYes, EC2 is dope. But so is DynamoDB, MSK, EKS, SQS, SNS, SES, IAM, KMS, etc. If all you use is EC2, then you need to manage the equivalent of a large number of services that are high quality and well run, and you'll need staff with expertise in running them well.
- cocoland2 3y agoTouches a chord this post.Systems like Kubernetes, Kafka are inherently complicated. My previous company got baremetal from AWS and installed k8s cluster on them. No offense to who architected it, we had multi country infra and made sense to take care of cost advantages on lower cloud costs using alternate providers We got a lot of critical infra running on them and then slowly there was tech-debt that would start accumulating. Clusters have to get updated , older DNS versions in k8s are slow, networking (Older Weave versions was bursting through the seams when the traffic exploded with many applications onboarded). SRE teams get overwhelmed, constant requests for adding PVC (Kafka & C* was on k8s) took a toll. Sanity prevailed in the end, there was decision to move to hosted PaaS infra, though I no longer work there, I just reminisced what we were going through. Though a "cloud-independent" solution will save pennies, it will definitely drown dollars in personnel costs and the uptime/SLA History repeats itself, because we don't learn from our mistakes (us or others)
- kyugkyugyuog 3y agoi hate k8s!!!!!!!!!!!
- qeternity 3y agoThere seem to be two camps: those who think k8s is a godsend, and those who think it's the devil incarnate. We fall into the former. We run Rancher across a couple of bare metal clusters and it's been mostly an amazing experience (ca. 3 years). The only issues we had were with Rancher specific bugs, but those have been resolved and for the most part our infra is pretty autonomous. We do all HA at the application layer, so local NVME as opposed to network storage. This means Patroni, Redis Sentinel/Cluster, etc. But it broadly just works. Maybe we're not big enough to bump into issues, but I couldn't imagine migrating to the labyrinth of vendor lockin masquerading as cloud services. What am I missing? Why do we have such a wildly different experience to others?
- SilverBirch 3y agoThis is a good example of people making questionable practical decisions based on good principles. Yes, in principle it'd be good if you were agnostic of your cloud provider. But there's a few issues with it. Firstly, you're going to work so much harder doing what you want to do on AWS whilst avoiding doing the things AWS wants you to do. They're not dumb! You don't want to be locked in? They really want you locked in and they're going to work very hard to make that happen. Secondly, the likelihood of you ever actually making the decision to leave is extremely low, so you're paying all these costs for what is at best a theoretically risk. And finally, even if you do everything perfectly and never depend on anything uniquely amazon.... leaving is still going to suck! It's still going to be a huge amount of work to migrate away!
- morelisp 3y agoI don't particularly like Kubernetes at all, and while I like Kafka it's definitely overkill for the kind of system discussed in the article. But I gotta say, 87%? What the hell? I had >98% uptime with the first Kafka tooling I ever built, and that was 3 nodes, shared machines with the ZKs, producers/consumers split across three cities, and processing 10x the traffic they're talking about; maintained by just me on medium-range Hetzner boxes. It feels like there's something deeper hiding here, more along the lines of "our developers really don't / can't care about how the software is operating in production."