6 ms·
I never understood the appeal of service meshes. Half of their reason to exist is covered by vanilla kubernetes, the rest is inter-node VPN (e.g. wireguard) and
by andsens 3y ago
I never understood the appeal of service meshes. Half of their reason to exist is covered by vanilla kubernetes, the rest is inter-node VPN (e.g. wireguard) and tracing (cilium hubble). Unless I’m missing something encrypting intra-node traffic is pretty silly.
K8S has service routing rules, network policies, access policies, and can be extended up the wazoo with whatever CNI you choose.
It’s similar to Helm, in that Helm puts a DSL (values.yaml) on top of a DSL (go templates) on top of a DSL (k8s yaml), just that it is routing, authentication, and encryption on top.. well, routing (service route keys), authentication (netpols), and encryption.
It boggles the mind!
- fragmede 3y agoLook, back in the day, things weren't encrypted, so you could listen in on your neighbor's phone calls, read their email, hack their bank accounts. Wireshark and etherdump and the most fun of all, driftnet. So, since then, everything has to be encrypted, lest someone hack there way to the family jewels. Never mind that the number of breaks to get there means there are usually bigger fish to fry. The important thing is to sprinkle magic encryption dust on everything because then we know it's Very Secure. (That's not to deride the fact that encryption is important, because it is, but sometimes it goes a bit far when there are other gaping holes that should be patched first.)
- anonzzzies 3y agoUsually, unless someone is really doing naive things, you will need to have access to a lot of almost physical things to sniff traffic. You almost need to physically have access to room where either the server or the client is, even with unencrypted traffic. People say; 'but they can sniff it at level3'; they sure can, IF they have actual access to level3 on a higher level than just using them for normal traffic. Hacked switch or router or so. Probably state actors can and do pull that off, but outside that, it's really not so easy to get to unencrypted traffic of just a random target. You still should encrypt things of course when you can, but you don't have to get quite that paranoid about it. All major hacks are 0-days (well, not updated Wordpress is not necessarily 0-day; a lot of 0-days are exploited months or years later), stolen credentials (social engineering usually), brute force password hacks or applications that are left open (root/root for mysql with 3306 open to the world). Those have nothing to do with (un)encrypted traffic.
- felixgallo 3y agoif you have the ability to execute code on a CPU, and that CPU is connected to a bus, and that bus is connected to a network card, you can sniff traffic. If you have data and business processes that include at least one entity A that lacks absolute trust at least one other entity B in your cluster, then the visible traffic of A by B is bad.
- anonzzzies 3y agoYes, but if you know that I run unencrypted traffic on my network and if I tell you that, you still won't be able to get to any of that if you cannot get into our network. Even if I tell you that I host at provider X and the traffic is unencrypted until it hits our webserver, you still won't be able to sniff any of it without getting very intimate with someone who has deeper access. Just hiring a machine at the same provider and putting the card in promiscuous mode is not going to get you anything from us.
- jabradoodle 3y agoIt's not just a specific actor targeting a specific entity though; it's any malicious dependency being ran in a privileged environment.
- anonzzzies 3y agoYes, that's true. But then you might have bigger issues I would say. But agreed. It's a good reason to make sure it's all closed off.
- nyrikki 3y agoLook at the default capabilities below, as a poster above mentioned NET_RAW and MKNOD are enabled by default. https://docs.docker.com/engine/reference/run/#runtime-privilege-and-linux-capabilities https://docs.docker.com/engine/reference/run/#runtime-privil... Unless you perfectly drop all privileges from every pod you are open to attack. Containers are not security contexts, they are namespaces, that require all actors that can launch a VM to actively drop privileges. This is an intentional design decision and not a bug.
- osigurdson 3y agoIf you didn't have helm, you would be writing your own regex scripts. I don't see how this would be better.
- numbsafari 3y agoIf you didn’t have helm, you’d be using one of the other, much better tools, and be happier for it.
- osigurdson 3y agoI generally meant, if you didn't have something like helm you would use regex. In any case, please elucidate, what tool do you like best?
- grncdr 3y agoNot the poster you're replying to, but when it comes to deploying your own applications: generating Kubernetes manifests with whatever language you're already using and feeding JSON to `kubectl apply -f -` can accomplish the same outcome with less effort. Helm is still useful for consuming 3rd party charts, but IMO it's status as the "default" is more due to inertia more than good design.
- osigurdson 3y agoIn my own experience I started off doing this because I wasn't ready to learn helm. However, after using helm once I didn't see the reason to do it in my own code any more.
- osigurdson 3y agoI get the sense for the responses (and downvoting) that I will eventually learn to dislike helm!
- rsanders 3y agoWhich ones would you recommend?
- bushbaba 3y agoService Meshes are something necessary for a small portion of Fortune 500s which have 1000s of microservices. Sure you could use load balancers but it becomes cost efficient to move towards a client-side load balancer. If you aren't a Google, Apple, Microsoft, ...etc scale company than a service mesh might be a tad overkill
- remram 3y agoIsn't kube-proxy already a client-side load-balancer?
- xyzzy_plugh 3y agoYou're close, but it's really when you have thousands of microservices using either shitty languages or shitty client RPC libraries where you can't easily perform client-side load balancing. There are plenty of languages and RPC frameworks where you can solve this without resorting to a service mesh. Practically, and to your point, service meshes solve an organizational problem, not a technical one.
- marcosdumay 3y agoI don't get this either. Doesn't the mesh become an scalability bottleneck just like load balancers? On that scale I'd expect people to use client-selected replicated services (like SMTP), and never something that centralizes connections (no matter where it's close to). You can always add observability at the endpoints. Unless your infrastructure is very unusual (like some random part of it costing millions of times more for no good reason, as on the cloud), this is not a big challenge; you add it to the frameworks your people use. (I mean, people don't go and design an entire service with whatever random components they pick, or do they?)
- crackez 3y agoWith Istio (envoy) you run a "sidecar" container in your pods which handles the "mesh" traffic, so it scales with the number of instances of your pods.
- 3y ago
- champtar 3y agoI agree that intra node encryption, if implemented by sidecars, is just wasting CPU cycles. Small note, unless it has changed recently, containerd default capabilities list includes CAP_NET_RAW, so hostNetwork=true pods can sniff all traffic.
- neya 3y agoI actually never understood the appeal of Kubernetes in the first place. I have production apps running on bare bones VMs serving millions of customers. Is this sort of complexity really necessary? At this point I would just consider serverless options. Sure, they would be a little more expensive, but that's a huge savings if we account for engineering teams' time.
- growse 3y agoCounter-take: I never understood the appeal of virtualizing the hardware. Is that complexity really necessary? Of course there's tradeoffs, but I think it's a specific perspective that says that Kubernetes is any more complex than virtualizing the hardware and scheduling multiple VMs across real hardware.
- pdimitar 3y agoYou didn't respond with any benefits of k8s though. What's the true value-add? Surely, coding an entire infrastructure in YAML is not it because that's horrific for anyone who wrote actually working software (one that didn't need 20+ commits of "try again" for a single feature to start working anyway).
- growse 3y agoI find it hard to believe that you've considered this for more than than two seconds and can't think of a single reason why k8s might be a good fit for someone's requirements. But here's one: It's a curated, extensible API that provides a decent abstraction over a heterogeneous collection of hardware. Nobody's done that before, and it's extraordinarily useful for being able to define intent. No-one forces you to use YAML, it's just a serialisation format. Once you have a bunch of components that implement this API, it becomes trivial to deploy pretty much any level of complexity of containerised application, without having to care too much about the actual exact details of how scheduling, networking, storage etc. is implemented. Even better, I can hit two different clusters configured in two completely different ways with the same manifest and get roughly the same result. It's the abstraction.
- kuhsaft 3y agoOne major usage of services meshes that I’ve come across is for the transparent L7 ALB. gRPC, which is now very common, uses long-running connections to multiplex messages. This breaks load-balancing because new gRPC calls, within a single connection, will not be distributed automatically to new endpoints. There is the option of DNS polling, but DNS caching can interfere. So, the L7 service mesh proxy is used to load balance the gRPC calls without modification of services. https://learn.microsoft.com/en-us/aspnet/core/grpc/loadbalancing?view=aspnetcore-8.0#why-load-balancing-is-important https://learn.microsoft.com/en-us/aspnet/core/grpc/loadbalan...
- pylua 3y agoI like that istio does mtls. It also helps with monitoring the requests.
- perrygeo 3y agoI've worked on several k8s clusters professionally but only a few that used a service mesh, Istio mainly. I'll give you the promise first, then the reality. The promise is that all of the inter-app communication channels are fully instrumented for you. Four things mainly 1) mTLS between the pods 2) Network resilience machinery (rate limiting, timeouts, retries, circuit breakers). 3) Fine grained traffic routing/splitting/shifting. And 4) telemetry with a huge ecosystem of integrated visualization apps. Arguably, in any reasonably large application, you're going to need all of these eventually. The core idea behind the service mesh is that you don't need to implement any of this yourself. And you certainly don't want to duplicate all of this in each of your dozens of microservices! The service mesh can do all the non-differentiated work. Your services can focus on their core competency. Nice story, right? In reality, it's a little different. Istio is a resource hog (I've evaluated Linkerd which is slightly less heavy weight but still). Rule of thumb: For every node with 8 CPUs, expect your service mesh to consume at least a CPU. If you're using smaller nodes on smaller clusters, the overhead is absurd. After setting up your k8s cluster + service mesh, you might not have room for your app. Second, as you mention, k8s has evolved. And much of this can be done, or even done better, in k8s directly. Or by using a thinner proxy layer to only do a handful of service-mesh-like tasks. Third, do you really need all that? Like I said, eventually you probably do if you get huge. But a service mesh seems like buying a gigantic family bus just in case you happen to have a few dozen kids.