6 ms·
Scaling Kubernetes to 7,500 nodes (2021)
- rmorey 4y agogood read. should probably get [2021] tag
- deleted 4y ago[deleted]
- sciurus 4y agoThis is from 2021 and was discussed then at https://news.ycombinator.com/item?id=25907312 https://news.ycombinator.com/item?id=25907312 I'm curious what they're doing now.
- dang 4y agoThanks! Macroexpanded: Scaling Kubernetes to 7,500 Nodes - https://news.ycombinator.com/item?id=25907312 https://news.ycombinator.com/item?id=25907312 - Jan 2021 (53 comments)
- MuffinFlavored 4y agoScaling it to 7,600 nodes (kidding)
- mdaniel 4y agoGiven that their first post in this vein was https://openai.com/research/scaling-kubernetes-to-2500-nodes https://openai.com/research/scaling-kubernetes-to-2500-nodes then one would expect it to be "Scaling it to 12,500 nodes" :-D $ kubectl get nodes I'm sorry, Dave, I can't do that
- MichaelMoser123 4y ago> I'm curious what they're doing now. building skynet, apparently. All powered by k8s!
- mrits 4y agoI'm not a huge fan of Kubernetes. However, I think there are some great use cases and undeniably some super intelligent people pushing it to amazing limits. However, after reading over this there are some serious red flags. I wonder if this team even understands what alternatives there are for scheduling at this scale or the real trade offs. It seems like an average choice at best and if I was paying the light bill I'd definitely object to going this route.
- Thaxll 4y agoThere is no alternative as far as I know. Which open source or private solution scale above 10k nodes and 100k apps ( pods )?
- emmp 4y agoOne of Nomad's major pitches is that it can scale larger than K8s. It's all over any comparison between the two.
- dharmab 4y agoAnd HashiCorp has a lot of success stories and experience making it work well. They can help you set up a continent-spanning cluster with 100k+ nodes.
- rco8786 4y agoMesos definitely does
- satvikpendem 4y agoIs Kubernetes simply BEAM but not on Erlang?
- b112 4y agoSuccess! Meanwhile, all 7500 nodes are, computationally, replaced by a 96 core, $10k server, in a dude's basement. With power to spare.
- intelVISA 4y agoBut I thought Good System Design involved reserializing the same data multiple times across the cloud(tm) and had a dedicated SRE and infra team - it's cheaper than one sys admin!
- jmillikin 4y agoYou'd generally want each of those 7500 machines be a full-sized server. No point running Kubernetes on tiny VMs, since its purpose is to provide bin-packed scheduling in a datacenter.
- electroly 4y agoThis isn't some dipshit enterprise running LOB software. This is OpenAI. These are all giant multi-GPU nodes getting slammed all day with machine learning jobs.
- mardifoufs 4y agoYeah, I'm sure openai could've trained gpt4 on a 10k$ machine.
- vvladymyrov 4y agoAlso they use Ray.io from Anyscale https://archive.ph/ZlMi5 https://archive.ph/ZlMi5
- osigurdson 4y ago>> Pods communicate directly with one another on their pod IP addresses with MPI via SSH It would be nice if someone could solve this problem in a more Kubernetes native way. I.e. here is a container, run it on N nodes using MPI- optimizing for the right NUMA node / GPU configurations. Perhaps even MPI itself needs an overhaul. Is a daemon really necessary within Kubernetes for example?
- antonchekhov 4y agoTo overcome the limitations on cluster size in Kubernetes, folks may want to look at the Armada Project ( https://armadaproject.io/ https://armadaproject.io/ ). Armada is a multi-Kubernetes cluster batch job scheduler, and is designed to address the following issues: A single Kubernetes cluster can not be scaled indefinitely, and managing very large Kubernetes clusters is challenging. Hence, Armada is a multi-cluster scheduler built on top of several Kubernetes clusters. Achieving very high throughput using the in-cluster storage backend, etcd, is challenging. Hence, queueing and scheduling is performed partly out-of-cluster using a specialized storage layer. Armada is designed primarily for ML, AI, and data analytics workloads, and to: - Manage compute clusters composed of tens of thousands of nodes in total. - Schedule a thousand or more pods per second, on average. - Enqueue tens of thousands of jobs over a few seconds. - Divide resources fairly between users. - Provide visibility for users and admins. - Ensure near-constant uptime. Armada is written in Go, using Apache Pulsar for eventing, Postgresql, and Redis. A web-based front-end (named "Lookout") provides easy end-user access to see the state of enqueued/running/failed jobs. A Kubernetes Operator to provide quick installation and deployment of Armada is in development. Source code is available at https://github.com/armadaproject/armada https://github.com/armadaproject/armada - we welcome contributors and user reports!