6 ms·
Running 1000 containers in Docker Swarm
- barhun 10y ago2k nodes, 100k containers https://blog.online.net/2016/07/29/docker-swarm-an-analysis-of-a-very-large-scale-container-system/ https://blog.online.net/2016/07/29/docker-swarm-an-analysis-...
- schmichael 10y ago5k nodes, 1 million containers, 5 minutes https://www.hashicorp.com/c1m/ https://www.hashicorp.com/c1m/ (Disclaimer: I'm on the Nomad team but wasn't at the time of the post)
- jacques_chester 10y agoI don't know much about Nomad and couldn't work out from the repo what the jobs were. If I guess correctly, it's an app using Redis. Is that correct? Disclosure: by coincidence of market forces, we're mortal enemies. Let's send christmas cards!
- schmichael 10y ago> it's an app using Redis. Is that correct? Yup! Repo could definitely be clearer, but here's the code: https://github.com/hashicorp/c1m/blob/master/schedbench/tests/nomad/classlogger/main.go https://github.com/hashicorp/c1m/blob/master/schedbench/test... Basically calls an increment in Redis and then blocks forever. > Disclosure: by coincidence of market forces, we're mortal enemies. Let's send christmas cards! Haha, hi mortal enemy! Christmas cards it is! If you're ever in Portland, OR I'll buy a beverage of your choice as well. :)
- jacques_chester 10y agoI extend the same offer in New York!
- jacques_chester 10y agoSince we're playing this game: 1.55k nodes, 250k containersed applications[0]. Mind you, it's hard to compare these as there's no real "cloud bench". For pure benchmark porn Nomad are the undisputed champs on their 1 million case. The Cloud Foundry scaling test was intended to show a system with fully service-configured, fully-routed apps, with varying app characteristics (memory and RPS). To further stress the system, thousands of apps crashing and are relaunched on a continuous basis. Cloud Foundry installations with >10k containers have been ordinary for a while now; the 250k thing was to ensure we had lots of headroom and shake out chokepoints in Diego. [0] https://content.pivotal.io/blog/250k-containers-in-production-a-real-test-for-the-real-world https://content.pivotal.io/blog/250k-containers-in-productio... Disclosure: I work for Pivotal, the majority donor of engineering on Cloud Foundry.
- titpetric 10y agoBig respect for your achievements. I guess at some point it just becomes the question of "where do i get a 1000 nodes" vs. "how do I run a 1000 containers". Or, more the justification for that amount of hardware - I mean, the one dream job which I would probably want is getting paid to cut out all the hardware use while keeping reliability/availability/functionality. Like these guys who cut their AWS bill by $1mil/year in about 3 months - https://segment.com/blog/the-million-dollar-eng-problem/ https://segment.com/blog/the-million-dollar-eng-problem/. The thing is that I'm not exactly sure where I'd fit in more - running this thing, or just fixing it for somebody else. I definitely know that I'm mostly dealing with pets and not cattle :)
- jacques_chester 10y agoWell, Cloud Foundry is deployed by BOSH. So you can, if you wish, use RackHD to deploy it to naked hardware (instead of OpenStack, GCP, AWS, Azure and I forget what else). Your apps will still be containerised, distributed and wired up the same way. There's always a point at which it makes engineering sense to flip the switch to doing it yourself. But that frontier is never static. We (plus our peers in the Cloud Foundry Foundation) and others in this space like Red Hat OpenShift are constantly pushing back the tipping point at which it makes economic sense to DIY. We already have very large customers with very large engineering teams, who've built platforms before. And they are switching because that effort no longer makes business sense. It's an expense they don't need for a platform they're the only maintainers of. One of our peers at IBM wrote about DIY[0]. We have our own much more markety-businessy whitepaper, with a very detailed case, on the same topic[1]. [0] https://hackernoon.com/stop-spending-engineering-effort-solving-problems-you-dont-have-8d18584f4d2a#.uykr1khc0 https://hackernoon.com/stop-spending-engineering-effort-solv... [1] https://content.pivotal.io/white-papers/the-upside-down-economics-of-building-your-own-platform https://content.pivotal.io/white-papers/the-upside-down-econ... Disclosure: I work for Pivotal, etc.
- eblanshey 10y agoDoes anyone know how easy it is to set up autoscaling with Docker Swarm running on Google Cloud or AWS? We're looking to get starting with Docker Swarm or Kubernetes soon, and are considering using Docker Swarm because of its simplicity and developer familiarity with Docker Compose (we use it for our dev environment). We just want to add nodes to a cluster as traffic spikes and subsides.
- KenCochrane 10y agoHave you looked at Docker for AWS yet? https://www.docker.com/aws https://www.docker.com/aws It will setup your swarm, which uses auto scaling groups for the worker nodes. You can then configure the auto scaling groups how ever you want, to scale based on your cloudwatch metrics, etc. There is also a Docker for GCP product in beta. https://beta.docker.com https://beta.docker.com but I don't know how auto scaling works for it. Disclaimer: I work at Docker on the Docker for AWS product.
- kingrolo 10y agoGoogle Container Engine supports cluster autoscaling to automatically add nodes with load. It's listed as a beta feature though. I've tried most of the Docker orchestration offerings and Container Engine seems by far the nicest. Swarm and Compose are really simple for getting up and running, but when we evaluated them there was still a missing piece required in that there was no neat way to do zero downtime deployments. There's a tool called Kompose to convert docker-compose config to kubernetes manifests (https://github.com/kubernetes-incubator/kompose https://github.com/kubernetes-incubator/kompose) although whilst it's nice to get you started we tend to maintain them separately now.
- acejam 10y agoI suggest looking into Amazon ECS. They have auto scaling features that can trigger based on container-level alerts and thresholds.
- hefeweizen 10y agoIn the context of Docker Swarm and Kubernetes, autoscaling refers to container level scaling ie. given a set of nodes, any autoscaling function would manage the number of containers that are currently running on these nodes. For instance/node level autoscaling (which is closer to what you need), I would recommend using the autoscaling features provided by AWS/Google Cloud.
- xchaotic 10y agoI always wonder, why not isolate on a process level, or even withing a single, multi-threaded app. Sure you can run some sort of web service on hundreds of docker containers or you can run a single, fast web server that scales?
- acejam 10y agoWhen that single web server goes down, it's not so "fast" anymore.
- undersuit 10y agoThat sounds like a fixable problem. I'm pretty sure Erlang programmers could give some tips. Why is worrying about a single web server going down more worrisome than some part of the Docker stack going down and causing the same issue?
- titpetric 10y agoActually, neither should be a problem if you have enough redundancy :) the hardest part of rolling your own infrastructure is testing mission critical systems (like databases) to be fault tolerant and at the same time reliable. Lots of great projects are out there that address some of these issues, but it takes a lot of attention to details (like transaction rates, ACID compliance, replication, etc.) to get it right. This is why a lot of developers which aren't in unicorn startups take advantage of technology which is available from giants like Amazon or Google, or specific problem-domain companies like CloudFlare for example. Netflix serves as a great example of a technology-driven company that is an inspiration to us, but there are so many others that really changed the way we approach problems - Tumblr, Etsy. But to stay on topic of netflix - I think their idea behind "chaos monkey" is great, and we're increasingly rolling out a (currently simple) docker swarm version of it - https://github.com/titpetric/docker-chaos-monkey https://github.com/titpetric/docker-chaos-monkey - the best way to eliminate worry is to test failure scenarios. As docker chaos monkey is designed to unpredictably "kill off" containers, your system gets the benefit of design to handle failures. It's one of those problems that you have to have a passion for however - it's like testing software. You're only testing software for the functionality and failures which you can predict, and I'm pretty sure that any of us can't predict all the ways in which software (or distributed systems) can fail. As such, it's a never ending occupation. :)
- officelineback 10y agoI'm still interested in how to merge features like AWS Autoscaling with Docker to right size the underlying infrastructure for the amount of container work going on.
- wwarren 10y agohttps://www.docker.com/aws https://www.docker.com/aws oughta be a good place to start
- peterwwillis 10y agoI would use IPv6 for the orchestration network, probably not touch the tcp/ip parameters except for port range (and open file descriptor), and break up the broadcast domain into smaller networks. It is not advisable to have thousands of machines on one broadcast domain, and it is a pain in the ass to troubleshoot, not to mention causes bigger headaches when one network problem affects all the nodes across the entire gigantic network.
- hefeweizen 10y agoSlight nitpick, but this articles deals with "Docker Swarm mode" [1], which is different from Docker Swarm [2]. [1] https://docs.docker.com/engine/swarm/ https://docs.docker.com/engine/swarm/ [2] https://github.com/docker/swarm https://github.com/docker/swarm [3] Difference between Docker Swarm and Swarm mode: http://stackoverflow.com/questions/40039031/what-is-the-difference-between-docker-swarm-and-swarm-mode http://stackoverflow.com/questions/40039031/what-is-the-diff...
- collyw 10y agoCan someone that needs to run workloads like this explain to me why this is needed? It sounds like over engineering for the sake of it. There are only so many apps at Facebook scale in the world.