6 ms·
good lord is this what modern microservices are like? How is it better have service A request to a proxy, which requests to another proxy, which requests to se
by phildenhoff 3y ago
good lord is this what modern microservices are like?
How is it better have service A request to a proxy, which requests to another proxy, which requests to service B? I get the security benefits of that, but the network architecture is boggling. How many PBs of data are sent each day for what could be a monolithic service?
Actually, to the end — companies that embrace microservices, which see the value in them. How do they manage network traffic for these kinds of K8s setups at scale? Surely they’re not using REST and HTTP. Is it as simple as protobufs over HTTP? Quic? Something else?
Edit: lots of great discussion below but I really meant how do they manage network TRAFFIC, not microservices in general :)
- throwup238 3y ago> How do they manage these kinds of K8s setups at scale? Our DevOps team starts off the monthly all hands meeting by leading a ritual during which they ceremoniously sacrifice an animal from the Fish and Wildlife Service's Threatened & Endangered Species list while the rest of the company chants: Exorcizamus te, omnis immundus spiritus omnis satanica potestas, omnis incursio infernalis adversarii, omnis legio, omnis congregatio et secta diabolica.
- zubairq 3y agohaha, IT staff doing mystic chants reminds me of this video I made over a decade ago: https://www.youtube.com/watch?v=8Sj3_NfDYeU https://www.youtube.com/watch?v=8Sj3_NfDYeU
- jacinda 3y agoI needed this laugh so much. Thank you.
- vaporary 3y ago"Getting a SCSI chain working is perfectly simple if you remember that there must be exactly three terminations: one on one end of the cable, one on the far end, and the goat, terminated over the SCSI chain with a silver-handled knife whilst burning *black* candles." -- Anthony DeBoer "SCSI is *not* magic. There are *fundamental* *technical* *reasons* why you have to sacrifice a young goat to your SCSI chain every now and then." -- John F. Woods
- jakupovic 3y agoI LOL-ed at this because I used to cry when no goats were around.
- spockz 3y agoThe data over the wire is the effect of having micro services and has little to do with service meshes. If anything, a service mesh could help by transparently enabling compression or upgrading to h3 without having to bake that in every app. Also, from a networking perspective, the proxies are hosted next to the instances so there is no difference there. We run 1800 services in our mesh. Most of them rest/http, a select few gRPC and graphql. We even have some soap/http services in it.
- slimsag 3y agowith 1800 services, do you feel like you are programming/architecting code in the same way a complex monolithic codebase might, working across them all? Or are the services just cogs in the machine, managed by individual devs / cogs in the machine? I imagine the latter and presume that's the main benefit of so many services, but genuinely curious as I've never experience that many services in an architecture
- bostik 3y agoI would say it's right about the middle of those two extremes. A well functioning service mesh (or even just a well maintained and discoverable ingress controller) is essentially invisible to the individual dev teams. Just think how modern stacks work from a frontend dev's perspective: team wants to use an additional feature, so they find the budgeted credit card, sign up to a random third-party provider, get their access token, and go. From the codebase standpoint, they merely added a new roundtrip to a random service and process the responses in their code. From the dev team's point of view, having the same feature available internally, behind "just another URL", makes no big difference. Maybe less politics around vendor spend and, with luck, easier integration with the remote service auth. Almost certainly less wrangling with compliance and legal. Whether that URL is provided by an ingress with a proper FQDN, or a service mesh entry with otherwise unresolvable name, is (and should be) irrelevant. Modern distributed systems have long since become too large for any single person to fully comprehend them through and through. There is no Grand Design[tm], they are all results of organic changes and evolution. Service discovery and routing can be architected. Individual services within the system can be architected. The complete system where hundreds or even thousands of services interact can not.
- de6u99er 3y agoThe trick is to know which code you want to maintain yourself and which stuff you want to get off the shelf. Going the off the shelf way, also means being stateless and using distributed transactions with all their pitfalls (eventual consistency, fire and forget, ...). Good architects will know how and when to do what E.g. start with a modulith instead of a monolith, to ease refactoring into micro-services once user count goes through the roof and vertical scaling won't do it any more.
- __turbobrew__ 3y ago> How do they manage these kinds of K8s setups at scale? You check in configs into the monorepo and there is tooling to continuously sync the configs with the actual state of the infrastructure. The advantage of microservices at scale is that team X breaking the build doesn’t affect team Y. This scale is probably not until you have 1000+ engineers however.
- pclmulqdq 3y agoTo be clear, that is the advantage of services. Microservices are a culture of taking that isolation as far as you can.
- jakupovic 3y agoYou do deployments as one used to do with big C++ programs also, you can delete everything and deploy again to make sure everything is new. The latter is preferred. Doing continuous deployments is very hard and not sure possible, at least I haven't seen it work well even though it's touted as fundamental k8s feature. K8s groups resources and presents a platform, how you use it it's up to you, and the management.
- pm90 3y ago> How is it better have service A request to a proxy, which requests to another proxy, which requests to service B? I get the security benefits of that, but the network architecture is boggling. How many PBs of data are sent each day for what could be a monolithic service? The tradeoff here is to decouple services in order to allow them to be developed somewhat independently of each other. Monoliths remove the overhead of network requests but they present their own challenges. You have a lot of implicit dependencies and feature development becomes complicated with changes having unexpected effects very far from the source. Ultimately engineering organizations need to decide the model that works best for them. Neither is inherently better, they’re solving different problems.
- deathanatos 3y ago> The tradeoff here is to decouple services in order to allow them to be developed somewhat independently of each other. You don't need a service mesh for that, though? Heck, you don't even need an ingress for service to service traffic. ingress-nginx does the job well, without being overly complex, and most importantly to me, logs when something is wrong, which I cannot say the same for Istio which I was fighting earlier this week where it was just happily RST'ing a connection and saying nothing about why it was deciding to do that.
- campbel 3y agoThere are a lot of good and bad reasons to adopt a mesh. Some of which might relate to your concerns. The things I like most about them, working in infrastructur: 1. I can have a unified set of metrics for all services, regardless of language/platform or how diligent the team is at instrumenting their apps. 2. I can guarantee zero trust with mTLS, without having to rely on application teams dealing with HTTPs or certificates. 3. I can implement automation around canary releases without much lift from dev teams. Other projects leverage these capabilities and do it for you as well. 4. I can get the equivalent of tcpdump for a pod pretty easily which I can use to help app teams debug issues. 5. I can improve app reliability with automatic retries and timeouts. Probably some other things as well... That said, it can be a big increase in complexity to your system the pains of which aren't always distributed to the folks getting the benefits.
- paulddraper 3y ago> A request to a proxy, which requests to another proxy, which requests to service B? I hesitate to tell you how many layers are between me an HN's servers right now.
- phildenhoff 3y agoRight but there are all those layers so that HNs one server can service your request and send a response. How much traffic would be generated if, for every request to HN, six other requests fire? And every time you comment a cascade of requests fire in HNs imaginary K8s cluster? It’s not that there is anything wrong with this, or that the tradeoff isn’t worth it… it’s just so much data flying back and forth over the wire.
- aetimmes 3y agoUltimately, in the age of massive cloud compute, the constraining resource in an organization is engineering hours, not CPU/memory/bandwidth. And even then, I've yet to encounter a system (outside of massively parallel MPI-based HPC) where saturating network pipes became the bottleneck for a system before CPU/memory utilization. Microservice architecture means there's lot of data flying around, but it keeps local resource utilization predictable.
- com 3y agoI thought that the constraining resource was money in all but the top 1000 or so businesses in the world with effectively 99%+ margins, mostly because they control their markets absolutely and can raise prices or change demand like the big three algorithmic advertising companies. Everybody else in the cloud goes to the wall without very careful cost control, which is by no means automated or low cost, either.
- eptcyka 3y agoMost often, the wire is virtual, its just copying bytes between processes.
- paulddraper 3y ago
- patcon 3y agoHeh wait until you hear how the cell signalling and any biological process in your body works :) It's an absolute clusterfuck of dependencies and side-effects. It's like space bar heating functionality[1] all the way down...! In my thinking, we might as well get used to the levels of indirection and complexity that AIs will be comfortable with. I suspect it will be more akin to what biological computation is "comfortable" with, and less like what our minds happen to prefer But I digress :) yes, microservices strike me as wiiild [1]: https://xkcd.com/1172/ https://xkcd.com/1172/
- 15457345234 3y agoYou're asking some really good questions here I find that microservices and this type of architecture have become a religion - you do it this way because you do it this way. You add another layer of complication because that's what you do now. You add this product because that's what you do now. Now you do it this way. Now you stop doing this thing and do this thing instead. It's all proclamations and a truly insane level of complexity and often a truly stunningly low level of performance achieved from some very powerful hardware because everything is behind at least twenty layers of abstraction and you're like, encrypting traffic which is just being passed between VMs which are on the same hardware, but because you can't guarantee that they're always on the same hardware you have to encrypt and use a proxy and... oh wow Watching it from the outside is a bit exhausting, it just seems to be so much churn and overhead.
- deleted 3y ago[deleted]
- threeseed 3y ago> for what could be a monolithic service On larger projects: a) It is sometimes impossible to run the entire application on a laptop. And so having a micro-service means you can quickly iterate and test before embedding it into the wider system. b) You will commonly run into conflicting transitive dependencies which you simply can't work around. Classic example being Spark on the JVM which brings in years old Hadoop libraries. c) The inter-relationships between component can become so complex that you really want to be able to use canary or green/blue deployment techniques to reduce risk.
- deathanatos 3y ago> Edit: lots of great discussion below but I really meant how do they manage network TRAFFIC, not microservices in general :) I don't really understand the question you're asking, but I think maybe the answer is that network pipes are just bigger than the scale most people are operating at? I don't think anything I've ever done has really had that many qps, and if it has, it is more likely to raise an eyebrow that says "who's spamming requests" more that "I guess we've made it to the big time". REST & protobufs are orthogonal. Empirically, literally nobody is doing REST, and most things are just ad hoc, poorly to not-at-all defined JSON/HTTP with a few HTTP verbs sprinkled in to make everyone feel good. It could be protobuf, too, if you like, but unless you have some truly gargantuan JSON, it really won't matter in the end. Compression will make up enough of the difference in size on the network. Some languages don't have to allocate the keys a billion times, too, though even in Python, it's a while before it starts to hurt. What I see more of is processes just inexplicably using gigabytes upon gigabytes of RAM, burning through whole years of CPU time for no particular reason before just dropping back to nominal levels like nothing happened, and dev teams that can't coherently understand the disconnect between just how much power a modern machine has, and what their design doc says their process is supposed to do (hint: something that shouldn't take that many resources). I like microservices, but there should be strong areas of responsibility to them. For most companies, I think that's ~2–3 services. At my current company, it's ~2 services + a database, with the rest being things like cron jobs or really small services that are just various glue or infra tooling.
- foota 3y agoTraffic within a data center is relatively cheap, is it not? Of course, I think managed network proxies tend to be expensive, but I think that mostly comes from the machine costs, not the network itself. You pay a bit of a tax for network serialization, but it's unlikely to be the bottleneck in most things ime. Plus, applications built on an external database (e.g., postgres not sqlite) will already have 1 hop. Going from 1 to 2 hops is less dramatic than going from no hops to 1. And I guess implicitly any service sending lots of data to the client already has 1 hop. Imo the only case where a network proxy would be egregious is an in memory database kind of workload where the response size is small relative to the size of the data accessed (maybe something like a custom analytics engine), but that's pretty niche.
- ozr 3y ago> Traffic within a data center is relatively cheap, is it not? Complexity is very, very expensive.
- manfre 3y ago> Traffic within a data center is relatively cheap, Most organizations will run their services spread across multiple data centers / AZs.
- foota 3y agoHm, wouldn't you want to set it up to route traffic preferentially to a nearby location?
- lukeschlather 3y agoI really hate this pattern in principle, but it isn't as bad as you're making it out to be. Most of the service meshes operate as sidecars, which is to say that if service A is calling service B through a mesh, there are two proxies in between the services, but proxy A is on the same machine as A, and proxy B is on the same machine as proxy B. So it is kind of offensive to have 3 network hops involved, but actually only one of those is over the wire, and in most cases I don't think the proxies are actually causing any measurable latency. (At least, whatever latency it's causing is smaller than the latency caused by the TLS encryption, which is necessary and the whole point.)
- deleted 3y ago[deleted]
- phildenhoff 3y agoThat makes sense! I didn't realise the proxies were sidecars, but that seems obvious in retrospect.
- dilyevsky 3y agoThey manage traffic with “service meshes” which includes anything from relatively” basic” distributed iptables/ipvs VIPs (kube-proxy) + dns to full-blown intercepting proxies a la istio/linkerd
- geraldwhen 3y agoIt’s all http with hundreds of milliseconds of latency per request.
- solatic 3y ago> How many PBs of data are sent each day for what could be a monolithic service? The companies that need service meshes couldn't possibly run everything in a single monolith. They already have several if not dozens of monoliths, each one typically coming into the architecture when a large enterprise acquires another company and its monolith, or simply different business units / product lines that don't talk to each other because of sheer organizational scale. It's called a service mesh, not a microservice mesh. First you get the benefits wrapping, securing, and monitoring each of your monoliths, then you reap more benefit when you start to break up each of those monoliths so common concerns can be addressed by common services and provide a unified experience to the enterprise's customers/users.
- deleted 3y ago[deleted]