6 ms·
Hi! Co-author here. We do keep the nodes running 24/7, so Kubernetes still provides the scheduling to decide which nodes are free or not at any given time. Gene
by benchess 6y ago
Hi! Co-author here. We do keep the nodes running 24/7, so Kubernetes still provides the scheduling to decide which nodes are free or not at any given time. Generally starting a container on a pre-warmed node is still much much faster than booting a VM. Also, some of our servers are bare-metal.
EDIT: Also don't discount the rest of the Kubernetes ecosystem. It's more than just a scheduler. It provides configuration, secrets management, healthchecks, self-healing, service discovery, ACLs... there are absolutely other ways to solve each of these things. But when starting from scratch there's a wide field of additional questions to answer, problems to solve.
- xorcist 6y agoIsn't Kubernetes a pretty lousy scheduler when it doesn't take this into consideration? There are a number of schedulers used in high performance computing that should be able to do a better job.
- AlphaSite 6y agoIf all you care about is node in use or not in use I think it’s fine. You don’t need anything complex from the scheduler.
- chubot 6y agoYeah exactly... This seems closer to an HPC problem, not a "cloud" problem. Related comment from 6 months ago about Kubernetes use cases: https://lobste.rs/s/kx1jj4/what_has_your_experience_with_kubernetes#c_qonwag https://lobste.rs/s/kx1jj4/what_has_your_experience_with_kub... Summary: scale has at least 2 different meanings. Scaling in resources doesn't really mean you need Kubernetes. Scaling in terms of workload diversity is a better use case for it. Kubernetes is basically a knockoff of Borg, but Borg is designed (or evolved) to run diverse services (search, maps, gmail, etc.; batch and low latency). Ironically most people who run their own Kube clusters don't seem to have much workload diversity. On the other hand, HPC is usually about scaling in terms of resources: running a few huge jobs on many nodes. A single job will occupy an entire node (and thousands of nodes), which is what's happening here. I've never used these HPC systems but it looks like they are starting to run on the cloud. Kubernetes may still have been a defensible choice for other reasons, but as someone who used Borg for a long time, it's weird what it's turned into. Sort of like protobufs now have a weird "reflection service". Huh? https://aws.amazon.com/blogs/publicsector/tag/htcondor/ https://aws.amazon.com/blogs/publicsector/tag/htcondor/ https://aws.amazon.com/marketplace/pp/Center-for-High-Throughput-Computing-HTCondor-899-/B073WHVRPR https://aws.amazon.com/marketplace/pp/Center-for-High-Throug...
- vergessenmir 6y agoIt maybe an HPC problem but I'm not sure the available solutions come close to k8s in terms of functionality and I'm not talking about scheduling. I used to work in HPC/Grid but it's been a while but I do remember Condor being clunky even though it had its uses. And the commercial grid offerings couldn't scale to almost 10k nodes back then (am not sure about now, or if they even exist anymore)
- toomuchtodo 6y agoCondor is clunky, but still in use in high energy physics, for example (LHC CMS detector data processing). For greenfield deployments, I would recommend Hashicorp's Nomad before Kubernetes or Condor if your per server container intent is ~1 (bare metal with a light hypervisor for orchestration), but still steer you to Kubernetes for microservices and web-based cookie cutter apps (I know many finance shops using Nomad, but Cloudflare uses it with Consul, so no hard and fast rules). Disclosure: Worked in HPC space managing a cluster for high energy physics. I also use (free version) Nomad for personal cluster workload scheduling.
- vergessenmir 6y agoI admit that Nomad is a fair middle ground due to its clean DSL and also because of the homogeneity of their workloads. The team at OpenAI used the k8s api to make extensions around multi-tenancy (across teams) to saturate available allocations, task specific scheduling modifications which were not supported by the k8s scheduler. I don't know if Nomad has this extensibility. Their plugins were around device plugins and tasks when I last looked at it.
- deleted 6y ago[deleted]
- jacobr1 6y agoExactly, we migrated to k8s not because we needed better scaling (ec2 auto scaling groups were working reasonably well for us) but because we kept inventing our own way to do rolling deploys or run scheduled jobs, and had a variety of ways to store secrets. On top of that developers were increasingly running their own containers with docker-compose to test services talking to each to each other. We migrated to k8s to A) have a way to standardize how to run containerized builds and get the benefits for "it works on my laptop" matching how it works in production (at least functionally) and B) a common set of patterns for managing deployed software. Resource scheduling only became of interest after we migrated when we realized the aggregation of our payloads allowed us to use things like spot instances without jeopardizing availability.
- stonogo 6y agoAre you starting from scratch? This architecture seems like a pretty standard HPC deployment with unnecessary containerization involved.
- hamandcheese 6y agoNot to me mention it’s a well known skillset that can more easily be hired for, as opposed to “come work on our crazy sauce job scheduler, you’ll love it!”
- dijit 6y agoI feel like we solved this problem over a decade ago (if you’re keeping machines warm anyway) with job brokers. Am I somehow mistaken?
- torbital 6y ago> self-healing, service discovery For a second I read that as self-discovery Damn kubernetes is some good shit
- yongjik 6y agoWell, considering how looking at kubernetes config makes me question the choices I have made in my life that led me into this moment, "self-discovery" is not too far off, I think.