7 ms·
Nomad, a cluster manager and scheduler
- porker 11y agoCongrats to Hashicorp, it's always exciting to see another release from you. Now I just have to find an excuse to use them... ;-)
- bluecmd 11y agoThat's... no small release. If it indeed has everything they claim that's extraordinary. What's the catch? Why haven't they made more noise about this?
- jedberg 11y agoI think this is their noise. :) Their big conference is today, so I expect a few more announcements.
- a-priori 11y agoAnyone able to give a compare/contrast here with other cluster management systems... Apache Mesos or Kubernetes, for example?
- nathankleyn 11y agoHashicorp themselves have published comparisons to Kubernetes, Mesos, et al on the Nomad site[1]. They look well written and generally not too biased. [1]: https://www.nomadproject.io/intro/vs/ https://www.nomadproject.io/intro/vs/
- covi 11y agohttps://www.nomadproject.io/intro/vs/mesos.html https://www.nomadproject.io/intro/vs/mesos.html I'm not sure I understand what is being said in this page. Is the Nomad scheduler centralized? If so, it has been demonstrated that distributed scheduling (e.g. Mesos) leads to better throughput and availability, while achieving a placement close to a centralized approach.
- Ateoto 11y agoKelsey Hightower plans on giving a talk on this tomorrow at HashiConf.
- sybhn 11y agoThey're building a nice little ecosystem...
- pm90 11y ago>Nomad is designed to be a global state, optimistically concurrent scheduler. Global state means schedulers get access to the entire state of the cluster when making decisions enabling richer constraints, job priorities, resource preemption, and faster placements. Can anyone shed light on how that's possible? I was under the impression that global state in distributed systems was not possible?
- josh2600 11y agohttp://research.google.com/pubs/pub41684.html http://research.google.com/pubs/pub41684.html Basically, there's a server that holds state for the cluster. When a scheduler attempts to load a job into the cluster, it grabs state from the aforementioned server, performs its job placement calculations, and then tries to submit its answer to the master state. Since there are many such schedulers, and the amount of time it takes to place a job is non-trivial, its possible that during the processing time to calculate placement, another scheduler might've consumed the requested resources. In this case, the first scheduler will provision whichever services are not in conflict for placement location, and perform a new calculation with the new state to place the remaining conflicted servers. The general idea is that most services don't have such a high affinity that the components need to be started all at the same time (a MapReduce job on 10,000 nodes can still run with 5,000 nodes while the second 5k are provisioning). The tradeoff here is that you have to manage conflicts, but the hope is that there are few enough that your optimism is rewarded. Does that make sense?
- derefr 11y agoSounds more like "greedy" than "optimistic" scheduling: the early jobs get the worms. In most scheduling systems, two 10,000-node jobs that each would saturate a cluster on their own will time-share if submitted together, with each job acting in practice more like 10,000 single-node jobs. The result is usually each job getting a probabilistic 50% share of the cluster while they're both running, and then whichever one runs longer saturating the cluster once the other ends. This scheduling system, meanwhile, would seem to just hand the 10,000 nodes over to job A, and then sleep job B until job A is done. Admittedly, in the case where the jobs aren't submitted at the same time, and job A has already grabbed and saturated the cluster, the two cases collapse together: job B must wait (unless you want to schedule processes rather than containers; then you just degrade the cluster's performance.) But for batch-processing applications, you'd usually schedule everything to start at once, precisely so that the scheduler could interleave the jobs.
- mapunk 11y agoSide note to Hashicorp devs: The Products section of your homepage is virtually unreadable on Windows/Chrome: https://i.imgur.com/8st8HQk.png https://i.imgur.com/8st8HQk.png
- nailer 11y agoHashi folk: this actually hit my own site a few weeks ago, learning from the experience: OS X renders fonts better, even on the same non-retina display, than Windows does. If you have something font-weight: 200 or less, it's fine on OS X, but it's completely unusable on Windows.
- mapunk 11y ago400 weight made it look a bit more clear for me.
- dedene 11y agoDoes anyone know how Nomad plays together with CoreOS?
- fidget 11y agoNo solution for persistent/statefull applications, which is a real disapointment, seeing as this is where I see orchestration systems currently breaking new ground and coming up with interesting solutions. The contraint systems also doesn't look too impressive; can I do the equiq of Marathon's GROUP_BY contraint (i.e. AZ GROUP_BY 2 -> ensure I have instances running on >=2 machines w/ different AZ values)? Also no maintenance primitives, but that's just me being in love with Mesos.
- maximegarcia 11y agoThat's always my first question. Let's have a pgsql "job", or a set of mongodb instance in cluster, backed by Docker. How do we orchestrate the data side ? Do we need to have special nodes where the data live in and set constraints for that ?
- binocarlos 11y agoI work for ClusterHQ and we make a tool called Flocker which aims to solve this problem.
- filearts 11y agoWhen the Hashicorp folk have a few minutes to breathe at their conf, I really hope that they could address the question of stateful apps. I imagine that users of Nomad also have persistent state to deal with and there must be a pattern that has emerged to solve this already?
- artursapek 11y agoHashicorp is all in on Golang.
- AYBABTME 11y agoGo is a pretty good choice of language to build stuff like that. It's easy to distribute, it's easy to write tools for, it's well known by a bunch of people in the domain.
- lobster_johnson 11y agoHaving looked at both Mesos (with its various frameworks) and Kubernetes, this immediately looks more attractive to me. No external dependencies (which arguably simplifies ops), competetive feature set, a nice job language, the fact that Docker is optional, polished documentation, etc. Having used some of Hashicorp's other products, I've grown to expect a level of quality and pragmatism that seems to present here, too. Not being JVM-based is a big plus in my book, too. How much production use has Nomad seen, I wonder?
- cpitman 11y agoThis one is interesting to me. I've been a big proponent of Hashicorp's other tooling, but this seems like an area that other projects are already addressing (and doing well in). Choice is great, but I think I would have preferred if they joined up with Kubernetes/Mesos/etc. Also, their messaging seems a little ingenious. Otto talks about how important it is to support microservice development and deployment, but Nomad lists as a con that Kubernetes has too many separately deployed and composed services. PS, I do work for Red Hat, so maybe I'm a little biased.
- amouat 11y agoDid you mean ingenious? Or disingenuous?
- cpitman 11y agoHa, you are correct, I meant disingenuous.
- rgarcia 11y agoAlso, their messaging seems a little disingenuous. Otto talks about how important it is to support microservice development and deployment, but Nomad lists as a con that Kubernetes has too many separately deployed and composed services. This is consistent with a (reasonable) belief that microservice architecture is an important design pattern to support, but may not be the best approach for all problems. From reading the docs, my sense is that Nomad takes the position that for a cluster scheduler, fewer moving parts leads to lower operational overhead, which outweighs any benefit that microservices may bring. E.g., it's more difficult to deploy a microservice platform like Nomad if the platform itself is deployed as a set of microservices.
- illamint 11y agoI think there's definitely a bootstrapping problem here: microservices are great if you have something like Kubernetes, Nomad, Mesos etc. on which to run and deploy them, but you have to run your platform on something and be able to bring it back up if it goes down and that's where I think Nomad might have the edge.
- cheeseprocedure 11y agoNomad appears to share some of Consul's internals, but it does not seem possible to use an existing Consul cluster as backing KV store/lock provider/etc. I'd like to understand why.
- EFruit 11y agoThis might stray a bit from the core topic, but does anyone have resources (preferably free) about the theory behind what Nomad covers? Batch/job/task Scheduling, etc.
- SEJeff 11y agoI'm genuinely not sure why they are trying to compete with the likes of Mesos or Kubernetes, or what they are really trying to achieve here. There is simply no way they'll build a community around nomad 1/2 as big as either of the aforementioned even if the software is really good.
- fapjacks 11y agoI'm sure to a fan that "really freaking loves Mesos" you can't imagine other software competing with Mesos, but Hashicorp writes excellent software generally. Serf and Consul have completely sold me on Nomad, and I haven't even used it yet. I think there will be similar draw for many people that have used Hashicorp's software before.
- SEJeff 11y agoNo it isn't about koolaid, it is about technology. Consul, Vagrant, and Vault are all excellent Hashicorp technologies I've used. They are great, and some of the best in the industry for their problem spaces. You've got Mesos scaling to 10,000+ physical node clusters today in production at the likes of companies like Apple and Twitter. You've got Kubernetes being adopted and developed by pretty much all of the open source heavyweights and it came from the experienced Google developed building... Google. Kubernetes ontop of Mesos is basically the holy grail in my personal opinion where you get the best Ops (mesos) story mixed with the best Dev (k8s) story. I guess we'll see how much Nomad takes off :) Hashicorp isn't a huge company, it seems like to me their best bet is keeping the focus relatively small so they can be the best at what they do. Even if Nomad is a huge hit and is amazing, it still seems kind of sad that they couldn't simply double down and help out with kubernetes. I do find it ironic that they talk about how nomad is for microservices and then make a dig at the several microservices that k8s is made up of.
- Rapzid 11y agoI'm sort of surprised to hear this view of the situation which is about the complete opposite of what I've had for the past few years. I have felt it's a shame that hashicorp have created consul, vault, a new raft implementation (which they got tons of flack for then etcd made THEIR OWN re-write which fixed a tons of issues this year..), serf, etc., and nobody has adopted or built on them. K8's current secrets solution is a bit underwhelming TBH, as an example. Kubernetes has support for 250 node clusters in their list of blockers: https://github.com/kubernetes/kubernetes/blob/master/docs/roadmap.md#blocking-features https://github.com/kubernetes/kubernetes/blob/master/docs/ro...
- larryweya 11y agoI've been using Mesos with Aurora in production for almost a year and while its very stable (zero downtime so far with loss of hosts a couple of times) the setup process was quite tedious and not something I'm looking forward to doing again. It also has a lot of moving parts - understand and setup zookeeper to get mesos up, understand and setup Aurora, use something for service discovery (I use AirBnB's Synapse for this). Plus I prefer to use tools I can choose to look under the hood of and perhaps make some contributions, which is a bit intimidating with both Mesos and Aurora (C/C++ and Scala). Because of this, I'm keen to try out Nomad mostly because of the promise of no-other-dependency/single binary plus the use of a single language across the stack - Golang