4 ms·
It’s an alternative, selectable by changing the proxy-mode flag on kube-proxy. If the iptables implementation is working for you, I wouldn’t necessarily jump to
by sisk 8y ago
It’s an alternative, selectable by changing the proxy-mode flag on kube-proxy. If the iptables implementation is working for you, I wouldn’t necessarily jump to it—though note that iptables has a big performance hit when you start getting to hundreds to thousands of overlay IPs and you’ll notice it if you manage a mid-to-large sized cluster. Certainly worth playing with (throw it on for a node or two for now), or worth the default for a new cluster, but like with any recently stable option, if it ain’t broke …
- tsenkov 8y agoLikely a dumb question (apologies in advance): can that in any way replace the need to spinning a native LoadBallancer on AWS?
- iooi 8y agoI don't think this a dumb question, ELBs on AWS are stupid expensive for what they do.
- skywhopper 8y agoI agree it's a reasonable question. And ELBs can be expensive at small scales. They can handle a huge amount of traffic with an exceptional level of reliability, but at small scales there are probably much better options available.
- sisk 8y agoNo need for apology: there’s a lot of _stuff_ when it comes to kubernetes networking. Using the IPVS mode instead of iptables won’t change anything with regards to external load balancers. The proxy mode of kube-proxy sets the mechanism by which cluster virtual IPs (i.e., the “Cluster IP” assigned to a service) are mapped to zero or more “real” pod IPs (typically assigned by an overlay network like Calico or Flannel). So, if I have a service that has been assigned a cluster IP of 10.10.10.10, kube-proxy uses iptables, IPVS, or a userspace proxy to intercept packets destined for that service and redirect them to one of the pods that fulfills that service’s label selector (maybe, for example, one of 10.20.0.1 or 10.20.0.2). Similarly with a LoadBalancer-type service, in addition to the above, your service will be assigned a nodePort for each port you've defined. kube-proxy will then add the proper rules—using the proxy-mode you've selected whether it's IPVS or not—to each node. This means that any traffic that hits any node on the selected nodePort will be redirected to one of the pods that fulfills the service label-selector. That's why if you look at the kubernetes-provisioned _real_ load balancer, you'll see it simply sends traffic to the selected nodePort on any of your nodes. IPVS, iptables, or the user space proxy (depending on the proxy-mode you've configured) is then responsible for sending the traffic to one of the right pods and ports from there. In other words, kube-proxy, by way of IPVS or iptables rules, makes sure that things destined for "services" (a not-real, totally virtual thing) end up at one of the right pods. It's a wholly inter-cluster routing concern. It's a lot but hopefully that makes sense.
- bogomipz 8y ago>"... note that iptables has a big performance hit when you start getting to hundreds to thousands of overlay IPs and you’ll notice it if you manage a mid-to-large sized cluster" Thanks for your reply. Could you elaborate on this performance issue? Since both iptables and IPVS are both in the kernel I would curious to hear why the issue doesn't exist between IPVS and overlays.
- sisk 8y agoStarting to fall outside of my domain so I encourage correction however one of the reasons, from what I understand, is that iptables resolves rules sequentially (top to bottom within the chain) versus IPVS which is a hash table lookup—the more rules, the slower iptables gets. The way that the loadbalancing works with the iptables implementation is rules with decreasing probability. That means that, worse case, the kernel with match and fall through one rule per replica. In other words, with 20 pods that match a service selector, 20 rules will be processed before the packet destination is rewritten to the right pod IP and port. Not a big impact when you’re dealing with a few rules but a tangible impact when you’re dealing with the amount of rules created by a good-sized kubernetes cluster.
- bogomipz 8y agoThanks for the detailed explanation. Yes in larger environments with thousands of iptables rules, using ipsets(hashing and constant time lookups) gets around this bottleneck) you are describing. I wonder if ipsets is not an option with Kube-proxy and iptables though?