3 ms·
To overcome the limitations on cluster size in Kubernetes, folks may want to look at the Armada Project ( https://armadaproject.io/ https://armadaproject.io/ ).
by antonchekhov 4y ago
To overcome the limitations on cluster size in Kubernetes, folks may want to look at the Armada Project ( https://armadaproject.io/ https://armadaproject.io/ ). Armada is a
multi-Kubernetes cluster batch job scheduler, and is designed to address the
following issues:
A single Kubernetes cluster can not be scaled indefinitely, and managing very large Kubernetes clusters is challenging. Hence, Armada is a multi-cluster
scheduler built on top of several Kubernetes clusters.
Achieving very high throughput using the in-cluster storage backend, etcd, is
challenging. Hence, queueing and scheduling is performed partly out-of-cluster
using a specialized storage layer.
Armada is designed primarily for ML, AI, and data analytics workloads, and to:
- Manage compute clusters composed of tens of thousands of nodes in total.
- Schedule a thousand or more pods per second, on average.
- Enqueue tens of thousands of jobs over a few seconds.
- Divide resources fairly between users.
- Provide visibility for users and admins.
- Ensure near-constant uptime.
Armada is written in Go, using Apache Pulsar for eventing, Postgresql, and Redis. A web-based front-end (named "Lookout") provides easy end-user access to see the state of enqueued/running/failed jobs. A Kubernetes Operator to provide quick installation and deployment of Armada is in development.
Source code is available at https://github.com/armadaproject/armada https://github.com/armadaproject/armada - we welcome
contributors and user reports!