11 ms·
Building the largest known Kubernetes cluster
- zoobab 10mo agoThe new mainframe.
- rvz 10mo ago> While we don’t yet officially support 130K nodes, we're very encouraged by these findings. If your workloads require this level of scale, reach out to us to discuss your specific needs Obviously this is a typical experiment at Google on running a K8s cluster at 130K nodes but if there is a company out their that "requires" this scale, I must question their architecture and their infrastructure costs. But of course someone will always request that they somehow need this sort of scale to run their enterprise app. But once again, let's remind the pre-revenue startups talking about scale before they hit PMF: Unless you are ready to donate tens of billions of dollars yearly, you do not need this. You are not Google.
- mlnj 10mo ago>You are not Google. It's literally Google coming out with this capability and how is the criticism still "You are not Google"
- Rastonbury 10mo agoThe criticism is at pre-PMF startups who believe they need something similar
- jcims 10mo agoI work for a mature public company that most people in the US have at least heard of. We're far from the largest in our industry and we run jobs with more than that almost every night. Not via k8s though.
- Tostino 10mo agoYou have jobs running on more than 130k different machines daily?? Are they cloud based VMs, or your own hardware? If cloud based, do you reprovision all of them daily and incur no cost when you are not running jobs? If it's your own hardware, what else do you do with it when not batch processing?
- jcims 10mo agoThey are provisioned on demand (cloud) and shut down when no longer needed.
- game_the0ry 10mo ago> You are not Google. 100% agree. People at my co are horny to adopt k8s. Really, tech leads want to put it on their resume ("resume driven development") and use a tool that was made to solve a particular problem we never had. The downside is now we now need to be proficient it at, know how to troubleshoot it, etc. It was sold to leadership as something that would make our lives easier but the exact opposite has happened.
- BruSwain 10mo agoI think k8s has a learning curve, absolutely, and there are absolutely cases where it can be unnecessary overhead. But I actually think those cases are pretty small. If you're running multiple apps, k8s is valuable. There is initial investment in learning the system, but its v-extensible, flexible, & portable. (Yes, every hyperscaler's implementation of k8s has its own nuance in certain places, but the core concept of k8s translates very well)
- game_the0ry 10mo agoWe must be terrible at implementation bc we have a had a prod outage and our DX is objectively worse. Its caused more problems and headaches for us.
- scottyah 10mo agoUse killercoda and get your CKA, I bet most of the confusion will be gone. I've basically started mandating it for newer folks on my team since it covers so many of the gaps that get created by people who try Just In Time learning on the systems. K9s is great for visual people who are used to vim.
- dilyevsky 10mo ago> You are not Google. You think they are just running it for fun? It's literally non-Google customers who wanted this as was explained in the article
- hazz99 10mo agoI’m sure this work is very impressive, but these QPS numbers don’t seem particularly high to me, at least compared to existing horizontally scalable service patterns. Why is it hard for the kube control plane to hit these numbers? For instance, postgres can hit this sort of QPS easily, afaik. It’s not distributed, but I’m sure Vitess could do something similar. The query patterns don’t seem particularly complex either. Not trying to be reductive - I’m sure there’s some complexity here I’m missing!
- phrotoma 10mo agoI am extremely Not A Database Person but I understand that the rationale for Kubernetes adopting etcd as its preferred data store was more about its distributed consistency features and less about query throughput. etcd is slower cause it's doing RAFT things and flushing stuff to disk. Projects like kine allow K8s users to swap sqlite or postgres in place of etcd which (I assume, please correct me otherwise) would deliver better throughput since those backends don't need to perform consenus operations. https://github.com/k3s-io/kine https://github.com/k3s-io/kine
- dijit 10mo agoYou might not be a database person, but you’re spot on. A well managed HA postgresql (active/passive) is going to run circles around etcd for kube controlplane operations. The caveat here is increased risk of downtime, and a much higher management overhead, which is why its not the default.
- Sayrus 10mo agoGKE uses Spanner as an etcd replacement.
- ZeroCool2u 10mo agoBut, and I'm honestly asking, you as a GKE user don't have to manage that spanner instance, right? So, you should in theory be able to just throw higher loads at it and spanner should be autoscaling?
- xyse53 10mo agoThey mention GCS fuse. We've had nothing but performance and stability problems with this. We treat it as a best effort alternative when native GCS access isn't possible.
- dijit 10mo agofuse based filesystems in general shouldn’t be treated as production ready in my experience. They’re wonderful for low volume, low performance and low reliability operations. (browsing, copying, integrating with legacy systems that do not permit native access), but beyond that they consume huge resources and do odd things when the backend is not in its most ideal state.
- thundergolfer 10mo agoAWS Lambda uses FUSE and that’s one of the largest prod systems in the world.
- dijit 10mo agoAn option exists, but they prefer you use the block storage API.
- thundergolfer 10mo agoNo, as in Lambda itself uses FUSE as an implementation detail of their container filesystem.
- dijit 10mo agoIt seems there were some major issues, but AWS has developed around them and optimised for its needs; (https://www.madebymikal.com/on-demand-container-loading-in-aws-lambda/ https://www.madebymikal.com/on-demand-container-loading-in-a...) Fair, but far from a common advice I’m willing to tell people (other CTOs) to do.
- 10mo ago
- blurrybird 10mo agoAWS and Anthropic did this back in July: https://aws.amazon.com/blogs/containers/amazon-eks-enables-ultra-scale-ai-ml-workloads-with-support-for-100k-nodes-per-cluster/ https://aws.amazon.com/blogs/containers/amazon-eks-enables-u...
- cowsandmilk 10mo agoThat is 100k vs 130k for Google’s new announcement. I can’t speak as to whether the additional 30k presented new challenges though.
- Cthulhu_ 10mo agoI want to believe that this is an order-of-magnitude kind of problem, that is, if 100K is fine then 500K is also fine. I only skimmed the article though, but I'm confident that it's more a physical hardware, time, space and electricity problem than a software / orchestration one; the article mentions that a cluster that size needs to be multi-datacenter already given the sheer power requirements (2700 watts for one GPU in a single node).
- belter 10mo ago130k nodes...cute...but can Google conquer the ultimate software engineering challenge they warn you about in CS school? A functional online signup flow?
- jasonvorhe 10mo agoFor what? Access to the control plane API?
- belter 10mo agoIn general... Try to sign up for their AI services...
- chrisandchris 10mo agoThe could team up with Microsoft, because their signup flow is fine but the login flow is badly broken.
- yanhangyhy 10mo agothere is a doc about how to do with 1M nodes: https://bchess.github.io/k8s-1m/#_why https://bchess.github.io/k8s-1m/#_why so i guess the title is not true?
- jakupovic 10mo agoDoing this at anything > 1k nodes is a pain in the butt. We decided to run many <100 nodes clusters rather than a few big ones.
- kvrty 10mo agoSame here. Non Kubernetes project originated control plane components start failing beyond a certain limit - your ingress controllers, service meshes etc. So I don't usually take node numbers from these benchmarks seriously for our kind of workloads. We run a bunch of sub-1k node clusters.
- liveoneggs 10mo agoSame. The control plane and various controllers just aren't up to the task.
- preisschild 10mo agoMeh, I've had had clusters with close to 1k nodes (w/ cilium as CNI) and didnt have major issues
- __turbobrew__ 10mo agoWhen I was involved about a year ago, cilium falls apart at around a few thousand nodes. One of the main issues of cilium is that the bpf maps scale with the number of nodes/pods in the cluster, so you get exponential memory growth as you add more nodes with the cilium agent on them. https://docs.cilium.io/en/stable/operations/performance/scalability/report/ https://docs.cilium.io/en/stable/operations/performance/scal...
- oasisaimlessly 10mo agoWouldn't that be quadratic rather than exponential?
- preisschild 10mo agoThats true and I definitely had to "tune" the bpf map limits, but it wasn't really that difficult to do.
- John-Tony 10mo ago[dead]
- blinding-streak 10mo agoImagine a Beowulf cluster of these
- sandGorgon 10mo agodoes anyone know the size at openai ? it used to run a 7500 node cluster back in 2021 https://openai.com/index/scaling-kubernetes-to-7500-nodes/ https://openai.com/index/scaling-kubernetes-to-7500-nodes/
- deleted 10mo ago[deleted]
- blamestross 10mo agoI worked in DHTs in grad school. I still double take that Google and other companies "computers dedicated to a task" numbers are missing 2 digits from what I expected. We have a lot of room left for expansion, we just have to relax centralized management expectations.
- Nextgrid 10mo agoK8S clusters on VMs strike me as odd. I see the appeal of K8s in dividing raw, stateful hardware to run multiple parallel workloads, but if you're dealing with stateless cloud VMs, why would you need K8S and its overhead when the VM hypervisor already gives you all that functionality? And if you insist anyway, run a few big VMs rather than many small ones, since K8s overhead is per-node.
- victorbjorklund 10mo agoBecause k8s gives you lots of other things out of the box like easy scaling of apps etc. Harder to do on VM:s where you would either have to dedicate one VM per app (might be a waste of resources) or you have to try and deploy and run multiple apps on multiple VM:s etc. (For the record I’m not a k8s fanatic. Most of the time a regular VM is better. But a VM isn’t = a kubernetes cluster).
- GauntletWizard 10mo agoThe reason to target k8s on cloud vms is that cloud VMs don't subdivide as easily or as cleanly. Managing them is a pain. K8s is an abstraction layer for that - Rather than building whole machine images for each product, you create lighter weight docker images (how light weight is a point of some contention), and you only have to install your logging, monitoring, and etc once. Your advice about bigger machines is spot on - K8s biggest problem is how relatively heavyweight the kublet is, with memory requirements of roughly half a gig. On a modern 128g server node that's a reasonable overhead, for small companies running a few workloads on 16g nodes it's a cost of doing business, but if you're running 8 or 4g nodes, it looks pretty grim for your utilization.
- nyrikki 10mo agoYou can run pods, with podman and avoid the entire k8s stack or even use minikube on a machine if you wanted to. Now that rootless is the default in k8s[0] the workflow is even more convenient and you can even use systemd with isolated users on the VM to provide more modularity and seporation. It really just depends on if you feel that you get value from the orchestration that full k8s offers. Note that on k8s or podman, you can get rid of most of the 'cost' of that virtualization for single placement and or long lived pods by simply sharing a emptyDir or volume shared between pod members. # Create Pod podman pod create --name pgdemo-pod # Create client podman run -dt -v pgdemo:/mnt --pod pgdemo-pod -e POSTGRES_PASSWORD=password --name client docker.io/ubuntu:25.04 # Unsafe hack to fix permissions in quick demo and install packages podman exec client /bin/bash -c 'chmod 0777 /mnt; apt update ; apt install -y postgresql-client' # Create postgres server podman run -dt -v pgdemo:/mnt --pod pgdemo-pod -e POSTGRES_PASSWORD=password --name pg docker.io/postgres:bookworm -c unix_socket_directories='/mnt,/var/run/postgresql/' # Invoke client using unix socket podman exec -it client /bin/bash -c "psql -U postgres -h /mnt" # Invoke client using localhost network podman exec -it client /bin/bash -c "psql -U postgres -h localhost" There is enough there for you to test to see that the performance is so close to native sharing unix sockets that way, that there is very little performance cost and a lot of security and workflow benefits to gain. As podman is daemonless, easily rootless, and on mac even allows you to ssh into the local linux vm with `podman machine ssh` you aren't stuck with the hidden abstractions of docker-desktop which hides that from you it has lots of value. Plus you can dump a k8s like yaml to use for the above with: podman kube generate pgdemo-pod So you can gain the advantages of k8s without the overhead of the cluster, and there are ways to launch those pods from systemd even from a local user that has zero sudo abilities etc... I am using it to validate that upstream containers don't have dial home by producing pcap files, and I would also typically run the above with no network on the pgsql host, so it doesn't have internet access. IMHO the confusion of k8s pods, being the minimal unit of deployment, with the fact that they are just a collection of containers with specific shared namespaces in the general form is missed. As Redhat gave podman to CNCF in 2024, I have shifted to it, so haven't seen if rancher can do the same. The point being is that you don't even need the complexity of minikube on VM's, you can use most of the workflow even for the traditional model. [0] https://kubernetes.io/blog/2025/04/25/userns-enabled-by-default/ https://kubernetes.io/blog/2025/04/25/userns-enabled-by-defa...
- supportengineer 10mo agoImagine a Beowulf cluster of these
- __turbobrew__ 10mo agoIt makes me sad that to get these scalability numbers requires some secret sauce on top of spanner, which no body else in the k8s community can benefit from. Etcd is the main bottleneck in upstream k8s and it seems like there is no real steam to build an upstream replacement for etcd/boltdb. I did poke around a while ago to see what interfaces that etcd has calling into boltdb, but the interface doesn’t seem super clean right now, so the first step in getting off boltdb would be creating a clean interface that could be implemented by another db.
- iwontberude 10mo agoFor those not aware, if you create too many resources you can easily use up all of the 8GB hard coded maximum size in etcd which causes a cluster failure. With compaction and maintenance this risk is mitigated somewhat but it just takes one misbehaving operator or integration (e.g. hundreds of thousands of dex session resources created for pingdom/crawlers) to mess everything up. Backups of etcd are critical. That dex example is why I stopped it for my IDP.
- scoodah 10mo agoThis is why I’ve always thought Tekton was a strange project. It feels inevitable that if you buy into Tekton CI/CD you will hit issues with etcd scaling due to the sheer number of resources you can wind up with.
- iwontberude 10mo agoYeah, quite unfortunate. But maybe there is hope. Apparently k3s uses Kine which is an etcd translation layer for relational databases and there is another project called Netsy which persists into s3 https://nadrama.com/netsy https://nadrama.com/netsy. Some interesting ideas. Hopefully native postgres support gets added since its so ubiquitous and performant.
- prescriptivist 10mo agoWhat boundaries does this 8GB etcd limit cut across? We've been using Tekton for years now but each pipeline exists in its own namespace and that namespace is deleted after each build. Presumably that kind of wholesale cleanup process keeps the DB size in check, because we've never had a problem with Etcd size... We have multiple hundreds of resources allocated for each build and do hundreds of builds a day. The current cluster has been doing this for a couple of years now.
- bhouston 10mo agoSounds like hell. But I do really dislike Kubernetes: https://benhouston3d.com/blog/why-i-left-kubernetes-for-google-cloud-run https://benhouston3d.com/blog/why-i-left-kubernetes-for-goog...
- jeffbee 10mo agoYou could remove all references to AI/ML topics from this article and it would remain just as interesting and informative. I really hate that we let marketing people cram the buzzword of the day into what should be a purely technical discussion.
- Heliodex 10mo agoView without needing to sign in: https://web.archive.org/web/20251124111136/https://cloud.google.com/blog/products/containers-kubernetes/how-we-built-a-130000-node-gke-cluster/ https://web.archive.org/web/20251124111136/https://cloud.goo...
- zkmon 10mo agoWhat business usecase requires a single cluster with thousands of pods? Wouldn't having multiple clusters, each hosting a few namespaces, be a better architecture?
- solatic 10mo agoThis. I may not work with AI training workflows, but I struggle to understand why they supposedly require launching a thousand pods per second to use GPUs that need to fundamentally be installed across different baremetal machines. Once the GPUs are on different machines, if there are 1k+ such machines, just start putting them on different Kubernetes clusters. Build a scheduling layer above the Kubernetes control plane to decide which Kubernetes cluster to schedule the pod onto. The whole thing stinks of, AI investors are throwing money at AI companies, so go to GCP and tell them to solve the problem at any price so that they can keep scaling without needing to build the scheduling layer above the Kubernetes control planes.
- zkmon 10mo agoYep, it's just saying "you should now launch 1000's pods in a single cluster, just because we said it makes sense, and please don't look at the costs, business sense and operational issues."
- moralestapia 10mo agoCute. I've done ~2 million (not k8s though, that trash would only slow me down).
- leo_e 10mo agoPapers like this are fascinating engineering, but dangerous marketing. They convince every Series A startup that they need a multi-region federated control plane for their 50 microservices. I spend half my time convincing my team not to emulate Google, because we don't have Google's scale problems—we have velocity problems. Complexity is an asset for Google (it's a moat), but a liability for the rest of us. I just want a cluster that doesn't require a dedicated ops team to upgrade.