5 ms·
I recently rebuilt my Kubernetes cluster running across three dedicated servers hosted by Hetzner and decided to document the process. It turned into a (so far)
by MathiasPius 3y ago
I recently rebuilt my Kubernetes cluster running across three dedicated servers hosted by Hetzner and decided to document the process. It turned into a (so far) 8-part series covering everything from bootstrapping and firewalls to setting up persistent storage with Ceph.
Part I: Talos on Hetzner
https://datavirke.dk/posts/bare-metal-kubernetes-part-1-talos-on-hetzner/ https://datavirke.dk/posts/bare-metal-kubernetes-part-1-talo...
Part II: Cilium CNI & Firewalls
https://datavirke.dk/posts/bare-metal-kubernetes-part-2-cilium-and-firewalls/ https://datavirke.dk/posts/bare-metal-kubernetes-part-2-cili...
Part III: Encrypted GitOps with FluxCD
https://datavirke.dk/posts/bare-metal-kubernetes-part-3-encrypted-gitops-with-fluxcd/ https://datavirke.dk/posts/bare-metal-kubernetes-part-3-encr...
Part IV: Ingress, DNS and Certificates
https://datavirke.dk/posts/bare-metal-kubernetes-part-4-ingress-dns-certificates/ https://datavirke.dk/posts/bare-metal-kubernetes-part-4-ingr...
Part V: Scaling Out
https://datavirke.dk/posts/bare-metal-kubernetes-part-5-scaling-out/ https://datavirke.dk/posts/bare-metal-kubernetes-part-5-scal...
Part VI: Persistent Storage with Rook Ceph
https://datavirke.dk/posts/bare-metal-kubernetes-part-6-persistent-storage-with-rook-ceph/ https://datavirke.dk/posts/bare-metal-kubernetes-part-6-pers...
Part VII: Private Registry with Harbor
https://datavirke.dk/posts/bare-metal-kubernetes-part-7-private-registry-with-harbor/ https://datavirke.dk/posts/bare-metal-kubernetes-part-7-priv...
Part VIII: Containerizing our Work Environment
https://datavirke.dk/posts/bare-metal-kubernetes-part-8-containerizing-our-work-environment/ https://datavirke.dk/posts/bare-metal-kubernetes-part-8-cont...
And of course, when it all falls apart: Bare-metal Kubernetes: First Incident
https://datavirke.dk/posts/bare-metal-kubernetes-first-incident/ https://datavirke.dk/posts/bare-metal-kubernetes-first-incid...
Source code repository (set up in Part III) for node configuration and deployed services is available at https://github.com/MathiasPius/kronform https://github.com/MathiasPius/kronform
While the documentation was initially intended more as a future reference for myself as well as a log of decisions made, and why I made them, I've received some really good feedback and ideas already, and figured it might be interesting to the hacker community :)
- AndrewKemendo 3y agoThank you for the amazing write up!
- baz00 3y agoAh man just looking at that list makes me glad for EKS. But thanks for the effort, I will read to learn more.
- MathiasPius 3y agoAbsolutely! If at all possible, go managed, preferably with a cloud provider that handles all the hard things for you like load balancing and so on. *Sometimes* however, you want or need full control, either for compliance or economic reasons, and that's what I set out to explore :)
- js4ever 3y agoAgreed, this is probably the best ad for managed k8s, this and horrors stories about self managed k8s clusters falling appart.
- msm_ 3y agoIf you ever want to have fun with setting up your own k8s, I recommend to start small. The author is already knowledgeable, so they probably knew from the start what they want, but a lot of this complexity is not essential. When I deployed my first kubernetes "cluster", I just spinned a single-node "cluster" using kubeadm (today k3s is an option too) and started deploying services (with no distributed storage - everything stored using hostPath). You only need to know kubernetes basics to do this. Then you probably want to configure CNI (I recommend flannel when starting, later cilium), spin an ingress controller (I recommend nginx or traefik), deploy cert-manager (this was hard for me when I started) and you can go a long way. With time I scaled up, decided to use GitOps, and deployed many more services (including my own registry - I started with docker's own, then migrated to Gitea. Harbor is too heavy for me). And of course over time you add monitoring, alerting etc - the fun never ends (but it's all optional, you should to decide when is the right time).
- smartbit 3y agoHave you tried KubeOne? Also with the benefits of machine-deployments. Works like a charm, didn’t go through your blogs, but KubeOne on Hetzner [0] seems easier than your deployment. And yes, also Open Source and German support available. [0] https://docs.kubermatic.com/kubeone/main/architecture/supported-providers/ https://docs.kubermatic.com/kubeone/main/architecture/suppor...
- MathiasPius 3y agoHetzner Cloud is officially supported, but that means setting up VPSs in Hetzner's Cloud offering, whereas this project was intended as a more or less independent pure bare-metal cluster. I see they offer Bare Metal support as well, but I haven't dived too deep into it. I haven't used KubeOne, but I have previously used Syself's https://github.com/syself/cluster-api-provider-hetzner https://github.com/syself/cluster-api-provider-hetzner which I believe works in a similar fashion. I think the approach is very interesting and plays right into the Kubernetes Operator playbook and its self-healing ambitions. That being said, the complexity of the approach, probably in trying to span and resolve inconsistencies across such a wide landscape of providers, caused me quite a bit of grief. I eventually abandoned this approach after having some operator somewhere consistently attempt and fail to spin up a secondary control plane VPS against my wishes. After poring over loads of documentation and half a dozen CRDs in an attempt to resolve it, I threw in my hat. Of course, Kubermatic is not Syself, and this was about a year ago, so it is entirely possible that both projects are absolutely superb solutions to the problem at this point.
- ralala 3y agoInteresting read. I have just setup a very similar cluster this week: 3 node bare metal cluster in a 10G mesh network. Decided for Debian, RKE2, Calico and Longhorn. Encryption is done using LUKS FDE. For Load Balancing I am using the HCloud Load Balancer (in TCP mode). At first I had some problems with the mesh network as the CNI would only bind to a single interface. Finally solved it using a bridge, veth and isolated ports.
- fireflash38 3y agoUsing containerd I assume? I've been trying to get RKE2 or k3s play nicely with CRI-O and it's been a long exercise in frustration.
- KyleSanderson 3y agowhich distro? it should just work out of the box.
- fireflash38 3y agoInitially Ubuntu 20.04, but I upgraded to 22.04. Finally got it working -- turns out a lot of things that reference `--cgroup-driver="systemd"` are doing it as if it were run in shell, which means that the quotes around "systemd" get removed by shell, and would lead to an error & ignored options. Nothing was showing whatsoever when using 20.04, so I wonder if there were some missing dependencies somewhere there... I'll probably write up everything I discovered at some point, there's a lot of pieces that you have to cobble together from pretty disparate sources (network plugins, config files (which!?), etc).
- cjr 3y agoGreat write up and what I especially enjoyed was how you kept the bits where you ran into the classic sort of issues, diagnosed them and fixed them. The flow felt very familiar to whenever I do anything dev-opsy. I’d be interested to read about how you might configure cluster auto scaling with bare metal machines. I noticed that the IP address of each node are kinda hard-coded into firewall and network policy rules, so that would have to be automated somehow. Similarly with automatically spawning a load-balancer from declaring a k8s Service. I realise these things are very cloud provider specific but would be interested to see if any folks are doing this with bare metal. For me, the ease of autoscaling is one of the primary benefits of k8s for my specific workload. I also just read about Sidero Omni [1] from the makers of Talos which looks like a Saas to install Talos/Kubernetes across any kind of hardware sourced from pretty much any provider — cloud VM, bare metal etc. Perhaps it could make the initial bootstrap phase and future upgrades to these parts a little easier? [1]: https://www.siderolabs.com/platform/saas-for-kubernetes/ https://www.siderolabs.com/platform/saas-for-kubernetes/
- MathiasPius 3y agoWhen it comes to load balancing, I think the hcloud-cloud-controller-manager[1] is probably your best bet, and although I haven't tested it, I'm sure it can be coerced into some kind of working configuration with the vSwitch/Cloud Network coupling, even if none of cluster nodes are actually Cloud-based. I haven't used Sidero Omni yet, but if it's as well architected as Talos is, I'm sure it's an excellent solution. It still leaves open the question of ordering and provisioning the servers themselves. For simpler use-cases it wouldn't be too difficult to hack together a script to interact with the Hetzner Robot API to achieve this goal, but if I wanted any level of robustness, and if you'll excuse the shameless plug, I think I'd write a custom operator in Rust using my hrobot-rs[2] library :) As far as the hard-coded IP addresses goes, I think I would simply move that one rule into a separate ClusterWideNetworkPolicy which is created per-node during onboarding and deleted again after. The hard-coded IP addresses are only used before the node is joined to the cluster, so technically the rule becomes obsoleted by the generic "remote-node" one immediately after joining the cluster.[3] [1] https://github.com/hetznercloud/hcloud-cloud-controller-manager https://github.com/hetznercloud/hcloud-cloud-controller-mana... [2] https://github.com/MathiasPius/hrobot-rs https://github.com/MathiasPius/hrobot-rs [3] https://github.com/MathiasPius/kronform/blob/main/manifests/infrastructure/cluster-policies/host-fw-control-plane.yaml#L79-L89C24 https://github.com/MathiasPius/kronform/blob/main/manifests/...