4 ms·
EKS has its own share of problems, and in several cases they can't be mitigated due to the limitations that EKS imposes. For example: - Service cluster IP rang
by symfrog 7y ago
EKS has its own share of problems, and in several cases they can't be mitigated due to the limitations that EKS imposes. For example:
- Service cluster IP range is not definable, making it difficult to integrate with an existing network topology [1]
- Limited choice of CNIs, e.g. Calico can not be used
- No accessible etcd snapshots for recovery
I tend to avoid EKS and prefer to create most clusters by using kubeadm as the foundation.
[1] https://github.com/aws/containers-roadmap/issues/216 https://github.com/aws/containers-roadmap/issues/216
- markbaikal 7y agoCalico can be used on EKS!? https://docs.aws.amazon.com/eks/latest/userguide/calico.html https://docs.aws.amazon.com/eks/latest/userguide/calico.html
- symfrog 7y agoCalico can only be used as the network policy engine on EKS, not as the CNI
- micah_chatt 7y agoEKS Engineer here. Calico policy can be used with the AWS VPC CNI, but you can remove the default CNI and install Calico or any other CNI plugin you’d like.
- symfrog 7y agoIn theory, you could replace the CNI on worker nodes, but is that something that is practically useful (when it can't be done on master nodes in EKS) and supported? How would the kube-apiserver, for example, communicate to the metrics-server if it is not connected to the Calico network?
- micah_chatt 7y agoYou are correct that the API server is only aware of the VPC network, and not any overlays. One solution to the metrics-server or other webhooks is to use host-networking mode so the API server can have connectivity.
- jrockway 7y agoIndeed, the networking layer that EKS forces upon you is quite strange. For example, you can only run 17 pods on a t3.medium instance because of how IP addresses are managed in your VPC (and because of various mandatory daemonsets that run pods on every machine, like aws-node). Most people will likely find out about this very low limit not in normal operations, but when one AZ blows up and pods are rescheduled on other AZs. You will hit the pod limit, things will be down and nothing new will schedule. And Amazon provides no meaningful monitoring for health at the control plane level; you need to bring all that yourself. Managing worker nodes is also a huge pain. They recently released a tool that manages some of it, but it doesn't work with older clusters. To do something simple like fix a vulnerability in the linux kernel, you have to make a new CloudFormation stack (involving cutting and pasting a ton of random stuff; the template ID, the AMI ID, etc.), edit several security groups, start the new nodes, make sure they join, drain the old nodes, delete the old stack, edit security groups. Upgrading the Kubernetes version in the cluster is a similar story, especially if you skipped an intermediate version. (That said, incidental changes like adding more nodes is easy. It's a lot rarer than upgrading Linux or the k8s version, however.) Their load balancers other than Classic don't work every well either. I would prefer to use a NLB for all incoming traffic (terminated by Envoy in the cluster which does the complicated routing) but that apparently results in kube-proxy not working right... and every time you change a set of pods whose selector matches the NLB service, it changes all the IP addresses in DNS, and traffic just stops arriving until the DNS TTL changes. It is very broken and the docs should say "never even think about using this" instead of "it's beta and here are some weird caveats that probably won't dissuade you from using it". Anyway, my overall impression of EKS is that Jeff Bezos walked into someone's office, said "you're all fired if we don't have some half-ass Kubernetes service in the next 3 months" and walked out. The result is EKS. It's wayyyyyy better than being locked into Amazon's crazy stuff (CloudFormation, etc.) but it's not as good as what other managed k8s providers offer. If I were to setup k8s again, I would just buy my own servers and self-host everything. Dealing with Things That Can Go Wrong With Servers is much easier than dealing with Things That Can Go Wrong With AWS. (But I already have a datacenter, can get the machines a 100Gbps connection to the Internet for free, and have a team of network engineers to tell me which CNI to use. If I didn't have that I would just use GCP.)
- micah_chatt 7y agoEKS Engineer here, thanks for the feedback. Service IP configurability is a very common ask, and as you’ve linked, is on our roadmap along with a slew of other control plane configuration options. You can delete the AWS VPC CNI DaemonSet and install any CNI plugin you’d like. EKS regularly backs up etcd and has automatic restore in the case of a failure. Manually restoring to an old snapshot would be quite disruptive. What is your use case, and what would be the interface you’d like to see?