8 ms·
Bare-Metal Kubernetes, Part I: Talos on Hetzner
- MathiasPius 3y agoI recently rebuilt my Kubernetes cluster running across three dedicated servers hosted by Hetzner and decided to document the process. It turned into a (so far) 8-part series covering everything from bootstrapping and firewalls to setting up persistent storage with Ceph. Part I: Talos on Hetzner https://datavirke.dk/posts/bare-metal-kubernetes-part-1-talos-on-hetzner/ https://datavirke.dk/posts/bare-metal-kubernetes-part-1-talo... Part II: Cilium CNI & Firewalls https://datavirke.dk/posts/bare-metal-kubernetes-part-2-cilium-and-firewalls/ https://datavirke.dk/posts/bare-metal-kubernetes-part-2-cili... Part III: Encrypted GitOps with FluxCD https://datavirke.dk/posts/bare-metal-kubernetes-part-3-encrypted-gitops-with-fluxcd/ https://datavirke.dk/posts/bare-metal-kubernetes-part-3-encr... Part IV: Ingress, DNS and Certificates https://datavirke.dk/posts/bare-metal-kubernetes-part-4-ingress-dns-certificates/ https://datavirke.dk/posts/bare-metal-kubernetes-part-4-ingr... Part V: Scaling Out https://datavirke.dk/posts/bare-metal-kubernetes-part-5-scaling-out/ https://datavirke.dk/posts/bare-metal-kubernetes-part-5-scal... Part VI: Persistent Storage with Rook Ceph https://datavirke.dk/posts/bare-metal-kubernetes-part-6-persistent-storage-with-rook-ceph/ https://datavirke.dk/posts/bare-metal-kubernetes-part-6-pers... Part VII: Private Registry with Harbor https://datavirke.dk/posts/bare-metal-kubernetes-part-7-private-registry-with-harbor/ https://datavirke.dk/posts/bare-metal-kubernetes-part-7-priv... Part VIII: Containerizing our Work Environment https://datavirke.dk/posts/bare-metal-kubernetes-part-8-containerizing-our-work-environment/ https://datavirke.dk/posts/bare-metal-kubernetes-part-8-cont... And of course, when it all falls apart: Bare-metal Kubernetes: First Incident https://datavirke.dk/posts/bare-metal-kubernetes-first-incident/ https://datavirke.dk/posts/bare-metal-kubernetes-first-incid... Source code repository (set up in Part III) for node configuration and deployed services is available at https://github.com/MathiasPius/kronform https://github.com/MathiasPius/kronform While the documentation was initially intended more as a future reference for myself as well as a log of decisions made, and why I made them, I've received some really good feedback and ideas already, and figured it might be interesting to the hacker community :)
- AndrewKemendo 3y agoThank you for the amazing write up!
- baz00 3y agoAh man just looking at that list makes me glad for EKS. But thanks for the effort, I will read to learn more.
- MathiasPius 3y agoAbsolutely! If at all possible, go managed, preferably with a cloud provider that handles all the hard things for you like load balancing and so on. *Sometimes* however, you want or need full control, either for compliance or economic reasons, and that's what I set out to explore :)
- js4ever 3y agoAgreed, this is probably the best ad for managed k8s, this and horrors stories about self managed k8s clusters falling appart.
- msm_ 3y agoIf you ever want to have fun with setting up your own k8s, I recommend to start small. The author is already knowledgeable, so they probably knew from the start what they want, but a lot of this complexity is not essential. When I deployed my first kubernetes "cluster", I just spinned a single-node "cluster" using kubeadm (today k3s is an option too) and started deploying services (with no distributed storage - everything stored using hostPath). You only need to know kubernetes basics to do this. Then you probably want to configure CNI (I recommend flannel when starting, later cilium), spin an ingress controller (I recommend nginx or traefik), deploy cert-manager (this was hard for me when I started) and you can go a long way. With time I scaled up, decided to use GitOps, and deployed many more services (including my own registry - I started with docker's own, then migrated to Gitea. Harbor is too heavy for me). And of course over time you add monitoring, alerting etc - the fun never ends (but it's all optional, you should to decide when is the right time).
- smartbit 3y ago
- wiktor-k 3y agoVery nice write-up! I wonder if it's possible to combine the custom ISO with cloud init [0] to automate the initial node installation? [0]: https://github.com/tech-otaku/hetzner-cloud-init https://github.com/tech-otaku/hetzner-cloud-init
- MathiasPius 3y agoI believe the recommended[1] way to deploy Talos to Hetzner Cloud (not bare metal) is to use the rescue system and Hashicorp Packer to upload the Talos ISO, deploying your VPS using this image, and then configuring Talos using the standard bootstrapping procedure. This post series is specifically aimed at deploying a pure-metal cluster. [1] https://www.talos.dev/v1.5/talos-guides/install/cloud-platforms/hetzner/ https://www.talos.dev/v1.5/talos-guides/install/cloud-platfo...
- wiktor-k 3y agoAh, I see. Thanks for the explanation!
- InvaderFizz 3y agoI'm going through you series now. Very well done. I thought I would mention that age is now built in to SOPS, thus needs no external dependencies and is faster and easier than gpg.
- MathiasPius 3y agoHave seen age pop up here and there, but haven't spent the cycles to see where it fits in yet, so I just went with what I knew. Will definitely take a look though, thanks!
- dhess 3y agoWhat performance numbers are you seeing on pods with Ceph PVs? e.g., what does `rados bench` give?
- MathiasPius 3y agoI haven't had an excuse to test it yet, but since it's only 6 OSDs across 3 nodes and all of them are spinning rust, I'd be surprised if performance was amazing. I'm definitely curious to find out though, so I'll run some tests and get back to you!
- MathiasPius 3y agoI rand rados benchmarks and it seems writes are about 74MB/s, whereas both random and sequential reads are running at about 130MB/s, which is about wire speed given the 1Gbit/s NICs. Complete results are here: https://gist.github.com/MathiasPius/cda8ae32ebab031deb0540542dc86f35 https://gist.github.com/MathiasPius/cda8ae32ebab031deb054054...
- dhess 3y agoThanks!
- wg0 3y agoI've come to the conclusion (after trying kops, kubespray, kubeadm, kubeone, GKE, EKS) that if you're looking for < 100 node cluster, docker swarm should suffice. Easier to setup, maintain and upgrade. Docker swarm is to Kubernetes what SQLite is to PostgreSQL. To some extent.
- vbezhenar 3y agoI didn’t try anything but kubeadm and it worked just fine for me for my 1 node cluster.
- wg0 3y agoBesides my local cluster of virtual box cluster, I have tried Kubernetes on three clouds with at least a dozen different installers/distributions and operational pain would be a factor going forward has always been my gut feeling. That's where the author also has following to say: >My conclusion at this point is that if you can afford it, both in terms of privacy/GDPR and dollarinos then managed is the way to go. And I agree. Kubernetes managed is also really hard for those of offering it and have to manage it for you behind the scenes.[0] [0]. https://blog.dave.tf/post/new-kubernetes/ https://blog.dave.tf/post/new-kubernetes/
- blowski 3y agoI agree in part - the features and simplicity of Docker Swarm are very appealing over k8s, but it also feels like so neglected that I'd be waiting every day for the EOL announcement.
- wg0 3y agoIt's built from another separate project called swarm-kit. So if it comes to that where it is abandoned, the forks would be out in the wild soon enough. I see more risk of docker engine as a whole pulling some terraform/elastic search licensing someday as investors get desperate to cash out.
- linuxdude314 3y ago
- xelxebar 3y agoSpeaking of k8s, anyone here know of ready-made solutions for getting XCode (i.e. xcodebuild) running in pods? As far as I'm aware, there are no good solutions for getting XCode running on Linux, so at the moment I'm just futzing about with a virtual-kubelet[0] implementation that spawns MacOS VMs. This works just fine, but the problem seems like such an obvious one that I expect there to be some existing solution(s) I just missed. [0]:https://github.com/virtual-kubelet/virtual-kubelet/ https://github.com/virtual-kubelet/virtual-kubelet/
- doctorpangloss 3y agoThere are no good ready made solutions. Someone has submitted patches to containerd and authored “rund” (d for darwin) to run HostProcess containers on macOS. The underlying problem is poorly familiarity with Kubernetes on Windows among Kubernetes maintainers and users. Windows is where all similar problems have been solved, but the journey is long.
- yjftsjthsd-h 3y agohttps://blog.darlinghq.org/2023/08/21/progress-report-q2-2023#the-future https://blog.darlinghq.org/2023/08/21/progress-report-q2-202... talks about running darling in flatpak, so it's not too much of a stretch to imagine it in a pod someday, but I don't think it's there today.
- lemper 3y agoI thought it was about talos the power9 system. intrigued by kubernetes on them.
- zkirill 3y agoMe too. That would be very cool and I'm surprised nobody is offering this as a service.
- mkagenius 3y agoFrom this, if people get the idea that they should get a Bare Metal on Hetzner and try. Don't. They will reject you probably, they are very picky. And if you are from a developing country like India, don't even think about it.
- mythz 3y agoThankfully we've never had the need for such complexity and are happy with our current GitHub Actions > Docker Compose > GCR > SSH solution [1] we're using to deploy 50+ Docker Containers. Requires no infrastructure dependencies, stateless deployment scripts checked into the same Repo as Project and after GitHub Organization is setup (4 secrets) and deployment server has Docker compose + nginx-proxy installed, deploying an App only requires 1 GitHub Action Secret, as such it doesn't get any simpler for us and we'll look to continue to use this approach for as long as we can. [1] https://servicestack.net/posts/kubernetes_not_required https://servicestack.net/posts/kubernetes_not_required
- seabrookmx 3y agoI used to do something similar at a previous company and this works well if you don't have to worry about scaling. YAGNI principal and all that. When you run hundreds of containers for different workloads, k8s bin packing and autoscaling (both on the pod and node level) tips the balance in my experience.
- mythz 3y agoYeah if we ever need to autoscale then I can see Kubernetes being useful, but I'd be surprised if this a problem most companies face. Even when working at StackOverflow (serving 1B+ pages, 55TB /mo [1]) did we need any autoscaling solution, it ran great on a handful of fixed servers. Although they were fairly beefy bare metal servers which I'd suspect would require significantly more VMs if it was to run on the Cloud. [1] https://stackexchange.com/performance https://stackexchange.com/performance
- swozey 3y agoI was a k8s contrib since 2015, version 1.1. I even worked at Rancher and Google Cloud. If you don't need absolutely granular control over a PAAS/SAAS (complex networking w/ circuit breaking yadda yadda, deep stack tracing, vms controlled by k8s (kubevirt etc), multi-tenancy in cpu or gpu) you don't need k8s and will absolutely flourish using a container solution like ECS. Use fargate and arm64 containers and you will save an absolute fortune. I dropped our AWS bill from $350k/mo to around $250k converting our largest apps to arm from x86. GKE is IMO the best k8s solution PAAS wise that exists, but quite frankly few companies need that much control and granularity in their infrastructure. My entire infrastructure now is AWS ECS and it autoscales and I literally never, ever, ever have had to troubleshoot it outside of my own configuration mishaps. I NEVER get on call alerts. I'm the Staff SRE at my corp.
- mulmen 3y agoJust finished reading part one and wow, what an excellently written and presented post. This is exactly the series I needed to get started with Kubernetes in earnest. It’s like it was written for me personally. Thanks for the submission MathiasPius!
- CoolCold 3y ago> Ceph is designed to host truly massive amounts of data, and generally becomes safer and more performant the more nodes and disks you have to spread your data across. I'm very pessimistic on CEPH usage in the scenario you have - may be I've missed it, but seen nothing about upgrading networking, as by default you gonna have 1Gbit on single interface used for public network/internal vSwitch. Even by your benchmarks, write test is 19 iops (block size is huge though) Max bandwidth (MB/sec): 92 Min bandwidth (MB/sec): 40 Average IOPS: 19 Stddev IOPS: 2.62722 Max IOPS: 23 Min IOPS: 10 while single HDD drive would give ~ 120 iops. single 3 years old NVMe datacenter edition, gives ~ 33000 iops with 4k block + fdatasync=1 CEPH would be very limiting factor in 1Gbit networking I believe - I'd put clear disclaimer on that for fellow sysadmins. P.S. The amount of work you done is huge and appreciated.
- sureglymop 3y agoHere's what I don't really get.. So, let's say you have three hosts and create your cluster. But now, you still need a reverse proxy or load balancer in front right? I mean not inside the cluster but to route requests to nodes of the cluster that are not currently down. So you could set up something like HAProxy on another host. But now you once again have a single point of failure. So do you replicate that part also and use DNS to make sure one of the reverse proxies is used? Maybe I'm just misunderstanding how it works but multiple nodes in a cluster still need some sort of central entry point right? So what is the correct way to do this.
- ralgozino 3y agoYou almost answered your own question. One common solution is to have 2 nodes with haproxy (or similar) sharing a virtual IP with keepalived that load balance de traffic to the control plane nodes and to the nodes where your ingress controller runs. There are other options, like running the haproxy in the control plane nodes.
- sureglymop 3y agoThank you, this was very helpful! I read up on keepalived and the used protocols now!
- MathiasPius 3y agoMy solution for this setup is having ingress controllers on all three nodes, and then specifying all three IPs in all DNS records. That way the end user will "load balance" based on the DNS randomization. Of course, if a node goes down, a third of the traffic will be lost, but with low TTLs and some planning, you can minimoze the impact of this.
- sureglymop 3y agoIt's an interesting approach. I did it a bit differently. I set up three Proxmox nodes on three hetzner servers. Then I deployed virtual routers. I then set up HAProxy and k3s nodes as LXC containers. What's nice about the whole setup is that a proxmox node can go down and it all still works. I will now set up keepalived as mentioned in the other reply so the HAProxies will also be fully HA. Proxmox also works well with zfs and backups. I set up the proxmox nodes manually and did the rest with terraform + ansible. One `terraform destroy` cleans up everything nicely. I wonder how the performance difference is between bare metal and k8s node in LXC.
- dave-at-koor 3y agoGreat post. We (Koor) have been going through something similar to create a demo environment for Rook-Ceph. In our case, we want to show different types of data storage (block, object, file) in a production-like system, albeit at the smaller end of scale. Our system is hosted at Hetzner on Ubuntu. KubeOne does the provisioning, backed by Terraform. We are using Calico for networking, and we have our own Rook operator. What would have made the Rook-Ceph experience better for you?