4 ms·
Ansible has scaled really poorly to thousands of hosts for us. Things we have run into: - Running a job against a single host will finish in 3 minutes... runni
by mdeeks 11y ago
Ansible has scaled really poorly to thousands of hosts for us. Things we have run into:
- Running a job against a single host will finish in 3 minutes... running that exact same job against thousands will take well over an hour and max out your machine.
- Running against more than around 3k hosts will somehow consume all 60GB of RAM and trigger the oom-killer
- CPU usage on the ansible runner is absurd for a large amount of hosts. We're currently using a c4.8xlarge (our biggest box) just to run deploy jobs and have them finish in a reasonable amount of time (10-15 minutes)
Slicing up our inventory into chunks and running them on different servers sucks big time and is pretty hacky. How do I combine the results? Can't do orchestration like "Run X on these roles first, then run Y on these roles when you're done".
Most likely what I'm going to do is have a single server execute ansible doing only the following in async (aka CPU friendly) mode:
- Upload a current copy of ansible to S3
- Upload the configs to the target machines with ONLY the secrets that role needs in plain text. (I'm not putting my vault secret on every box!)
- Have the servers pull it down and execute in --connection=local mode.
- Wait until each remote finishes
All that said, I LOVE LOVE writing stuff in Ansible. It is so easy to read, follow, and understand. I picked up most of it in a day or two just by reading their "Best Practices" page. Getting it to work at scale hurts though :(
- deleted 11y ago[deleted]
- harshreality 11y ago> - Running against more than around 3k hosts will somehow consume all 60GB of RAM and trigger the oom-killer Have you looked at Salt (salt stack)?
- pyre 11y ago> It is so easy to read, follow, and understand I never liked that the variable namespace is global, so there isn't anyway for a module to be self-contained. If you execute Module 1, and then Module 2, Module 1 can set a variable that inadvertently affects Module 2. The "recommended" way around this is to prefix all variable names, but this becomes unwieldy very quickly as your variable names grow in size. The "nicer" way to do it would be to set have a dict/hash of variables, but that makes top-level overrides difficult because there is no way to override "hash_name.variable_name" you basically have to override the whole "hash_name" variable or nothing at all. I found it difficult sometimes to reason about how these variables would work, and if I needed to add something to defaults.yml or variables.yml in a module.
- LukeHoersten 11y agoAgreed. In fact variables are scoped by "play" not by "role" which is almost weirder. For those who don't know the organizational structure from largest to smallest is playbook->plays->roles->modules/tasks. Variable scope is at the "play" level. More than that, roles are meant to be distributable components on the Ansible Galaxy service they run. Galaxy gets almost no use because modularity and reusability is broken by having no idea how the role author namespaced their variables. Collisions happen all the time. Why we must manually manage scope with naming conventions when the computer can do it automatically with scope is beyond me. This is Ansible's biggest downside in my opinion. I've talked with the core devs about it on IRC and they (bcoca) agreed but thought it too late to make such a pervasive change as introducing role-level scope. As some of the other posts have mentioned, I still love Ansible despite this shortcoming.
- andrewvc 11y agoCan't agree enough here. The core deva act like it isn't a problem. The lack of encapsulation makes reusability impossible
- krakensden 11y agoI remember being really down on Ansible Galaxy, and whenever someone tried to use a community role, I'd pull it down, audit it, and make them fork it, because it was inevitably dangerous, not thought out, and with no tests. Now I'm back in a Chef shop, and Chef has a ton of tooling for re-usability, has put loads of thought and effort into the problem, and there are tons of cookbooks, many maintained by Chef, Inc. The problem isn't really any better though- the official and officially blessed cookbooks are still terrible, broken and unsafe in all sorts of obvious ways. I'd rant more, but I'd have to get specific and mean about actual people. It's the problem space, honestly. It all depends on the guardrails your workflow provides, and there's never enough in common. Anyway, there's a reason everyone loves golden-images.
- LukeHoersten 11y ago
- movedx 11y ago> Running a job against a single host will finish in 3 minutes... running that exact same job against thousands will take well over an hour and max out your machine. Have you used the variety of options to throttle this down? Running against 3,000 hosts at once is, in my opinion, absolutely crazy. And almost an contradiction to: > Slicing up our inventory into chunks and running them on different servers sucks big time and is pretty hacky. How do I combine the results? ... How do YOU combine them all? All results from all 3,000 hosts.? And it's not "hacky" at all, it's what you should be doing as a systems engineer/architect in the first place. No one network should be 3,000 hosts in size in a flat network structure; that should be partitioned up into logical arrangements for easier management (which you currently don't have, as per your own points), not to mention security. As for combining the results. Have you considered writing a simple module for Ansible which sits locally and it called after every task is complete? It's easy to work with and you can push the results into ElasticSearch. Ansible isn't a magic bullet, it's a tool on which to build, so build on it :-)