27 ms·
Why Puppet, Chef, Ansible aren't good enough
- geerlingguy 13y agoSo, basically, replace yum, apt, etc. with a 'stateless package management system'. That seems to be the gist of the argument. Puppet, Chef and Ansible (he left out Salt and cfengine!) have little to do with the actual post, and are only mentioned briefly in the intro. They would all still be relevant with this new packaging system. For some reason, this came to mind: https://xkcd.com/927/ https://xkcd.com/927/
- NyxWulf 13y agoYeah that xkcd came to my mind as well. I'm amazed at the timeless truths he captures in those comics. Creating another version which may be better isn't the hard part. Gaining consensus and getting people to give up the other ones is the hard part.
- talkingquickly 13y agoAnd in a way which so many people relate to. As soon as I saw "xkcd" I knew which one was going to be linked to...
- saraid216 13y agoI still want to do a thing where response codes are returned the way they are for HTTP. Query: Why are you screwing around? Response: XKCD 303.
- sdegutis 13y agoAs existing solutions get more painful, people quickly adopt less painful (i.e. generally better) solutions. So it's tied just as much to how much the old ways suck as to how good the new ways are.
- bryanlarsen 13y agoTrue, but: `sudo apt-get install nginx` just works. Perhaps they're doing it "wrong", but there are thousands of people who are making sure it just works. I have some of the problems described in the article, but it only happens when I can't use the package manager. It happens when I have to compile from scratch, or move things around or mess with config files, et cetera. For me, phoenix servers and docker are the solutions. Maybe they're not as pretty as what he describes, but there is a solution that works.
- josegonzalez 13y ago`sudo apt-get install nginx` does not just work if: - The repo is down - External network doesn't work - You are missing a dependency not in your apt-cache - It conflicts with another package due to a dependency All of which are possible and happen.
- rmc 13y agoThere are lots of people ensuring that doesn't happen. Sure it's possible, but it's unlikely. There are loads of debian/ubuntu apt mirrors. apt-get (or aptitude) downloads dependencies, package maintainers ensure that there aren't those conflicts.
- haberman 13y agoThis reminds me of a particularly devious C preprocessor trick: #define if(x) if ((x) && (rand() < RAND_MAX * 0.99)) Now your conditionals work correctly 99% of the time. Sure it's possible for them to fail, but unlikely. Now you might object that C if() statements are far more commonly executed than "apt-get install". This is true, but to account for this you can adjust "0.99" above accordingly. The point is that there is a huge difference between something that is strongly reliable and something that is not. Things that are unreliable, even if failure is unlikely, lead to an endless demand for SysAdmin-like babysitting. A ticket comes in because something is broken, the SysAdmin investigates and found that 1 out of 100 things that can fail but usually doesn't has in fact failed. They re-run some command, the process is unstuck. They close the ticket with "cron job was stuck, kicked it and it's succeeding again." Then go back to their lives and wait for the next "unlikely but possible" failure. Some of these failures can't be avoided. Hardware will always fail eventually. But we should never accept sporadic failure in software if we can reasonably build something more reliable. Self-healing systems and transient-failure-tolerant abstractions are a much better way to design software.
- berkay 13y agoAuthor's point is that we are focusing on the wrong place (puppet, chef, etc.), hence potentially making the problem worse by attempting to deal with the symptoms instead of addressing the root cause. Puppet/Chef etc. may indeed still be relevant and potentially even more widely used as it would become much simpler to develop recipes, etc.
- geerlingguy 13y agoI think the takeaway is that package management is ripe for improvement, but CM tools (and more flexible tools like Ansible) do so many more things besides automating package management that I question the author's mention of them for anything besides a hook to get more pageviews. This article has little to do with config management. A better title would be "apt, yum and brew aren't good enough; we can do better".
- girvo 13y agoNo, NixOS and NixOps tackle both, and you can't have the latter without the former. They are all intertwined, hence talking about both in the article :)
- cwp 13y agoNo. Well, yes, replace yum, apt, etc. But once you have a functional package management system, you don't need Puppet, Chef or Ansible, because the same stateless configuration language can be used to describe cluster configurations as well as packages. So build a provisioning tool based on that, instead. That provisioning tool is called NixOps. The article links to it, but doesn't really go into detail about NixOps as a replacement for Puppet et al.
- lmm 13y agoThat's just as true without Nix. I've worked somewhere that applied changes to its clusters by building debs; all you need is something that regularly executes apt-get and you're golden. (Of course, whether something's a good language for expressing particular kinds of tasks is another question)
- jes5199 13y agoam I going to switch distros just to use a different provisioning tool?
- sparkie 13y agoYes... Eventually. The Nix model is the only sane model going into the future where complexity will continue to increase. Nobody is forcing you to switch distros now, but hinting that you'll probably want to look into this, because the current models might well be on their way to obsolescence.
- westurner 13y ago* Unique PREFIX; cool. Where do I get signed labels and checksums? * I fail to see how baking configuration into packages can a) reduce complexity; b) obviate the need for configuration management. * How do you diff filesystem images when there are unique PREFIXes in the paths? A salt module would be fun to play around with. How likely am I to volunteer to troubleshoot your unique statically-linked packages?
- deleted 13y ago
- kzahel 13y agoI think I find myself in a minority that thinks "sudo apt-get install nginx" is much simpler and who doesn't care about edge cases. If there's an edge case, something is wrong with my machine and it should die.
- liveoneggs 13y agoworks great but doesn't scale well, is all. rdist was a good method in the old days which scaled a little better than one-off, but we're in the pull vs push world now.
- spindritf 13y agoHow does it not scale? Apt scaled to however many millions of machines run Debian and its derivatives. If you have a really large deployment, you can set up your own repository, or a mirror of existing repositories. If that is still not enough, you're Facebook.
- sanswork 13y agoTyping apt get install across a bunch of machines doesn't scale for the person doing it. The repository will scale.
- byroot 13y agoI think what he meant by do not scale well is that sshing into 1000 of your servers to manually run apt-get just don't work. apt-get is fine, it's just that at some point you need an automation tool to trigger it. A lot of chef recipes rely on the underlying package manager, that's fine.
- pdpi 13y agoIt's not Apt that doesn't scale, it's using apt-get manually on a large number of machines that's the problem. By the time you're done putting all your apt-gets and your config files and whatnot in some shell script to automate it all away, you've reinvented a poor clone of puppet/ansible/chef.
- 13y ago
- reitzensteinm 13y agoPersonally, I use Fabric for automation, and it's got all the problems the author says; if you get the machine into an unknown state, you're better off just wiping it and starting fresh. However, with the rise of virtual machines, that's a trivial operation in many cases. Even on bare metal hardware it's not a big deal, as long as you can tolerate a box disappearing for an hour (and if you can't, your architecture is a ticking time bomb). In fact, starting from a clean slate each time basically makes Fabric scripts a declarative description of the system state at rest... if you squint hard enough.
- cwp 13y agoI've used Fabric this way as well, and found the same thing. If you start from a known-good state and apply a deterministic transformation to it, you've got another known-good state. The problem is that running commands via SSH isn't really a deterministic process. It's fairly predictable, but it relies on a lot of inputs that can change unexpectedly. So it mostly works, but building a server has to be cheap and easy, because you throw them away a lot. I think my ideal system would be sort of a cross between this and NixOS. We build servers with NixOS-style stateless declarative configuration. But then instead of using NixOS mechanisms for functionally/deterministically modifying the running system, we treat running servers as immutable and disposable and build new ones whenever we need to change something. That seems like it would give maximum maintainability and robustness.
- copergi 13y agoDoesn't NixOps already give you exactly what you are talking about?
- cwp 13y agoYeah, it's pretty close. It seems to want to modify instances in place, though. On NixOS, that's a lot safer than other distributions, but I think I'd still prefer to handle application deployment the same way as scaling and fault management. That is, bring up new servers with the new version, gradually move traffic over to them, then shut down the old servers. I'm also wondering if NixOps can build AMIs the same way Aminator does, by mounting a volume on a builder instance, writing files to the volume, then snapshotting and making an AMI based on the snapshot. (In contrast to snapshotting a running system.) That would be truly awesome. I've only been playing with NixOps a short while, maybe there's already a way to do these things, or they're easy to implement on top of NixOps. Looks very cool, though.
- onalark 13y agoIf you've ever needed version X.Y of Package Z on a system, and all of its underlying dependencies, or newer versions than what your operating system supports, you know exactly what Domen is talking about. It's a good write-up. The idea of a stateless, functional, package management system is really important in places like scientific computing, where we have many pieces of software, relatively little funding to improve the quality of the software, and still need to ensure that all components can be built and easily swapped for each other. The HashDist developers (still in early beta: https://github.com/hashdist/hashdist https://github.com/hashdist/hashdist ) inherited a few ideas from Nix, including the idea of prefix builds. The thing about HashDist is that you can actually install it in userspace over any UNIXy system (for now, Cygwin, OS X, and Linux), and get the exact software configuration that somebody else was using across a different architecture.
- pmahoney 13y ago> The thing about HashDist is that you can actually install it in userspace over any UNIXy system (for now, Cygwin, OS X, and Linux) I've been looking into Nix (the package manager) recently, and it also can be installed on OS X and various Linux distros (I've seen Cygwin mentioned, but I don't think this is well supported). If you're familiar with both, what does HashDist provide over Nix?
- onalark 13y agoSure! As a disclaimer, I'm one of the developers of HashDist. There's a longer explanation here: http://hashdist.readthedocs.org/en/stable/faq.html http://hashdist.readthedocs.org/en/stable/faq.html but the crux is that Nix enforces purity, whereas HashDist simply defaults to pure. As an example, if you'd like to build a Nix system, you're going to need to compile or download an independent libc. HashDist is capable of bootstrapping an independent libc as well, but we default to using the system libc. We choose to seat our "pure" system on top of the system compilers and libc, but by default, install our own Python and other high-level libraries. We also can integrate with other package managers, and there's an example branch in our repository right now showing integration with Homebrew.
- liveoneggs 13y agoauthor has a shallow or non-existent understanding of making pkgs for an operating system, doing system's administration, or the automation tools mentioned.
- berkay 13y agoIf you're going to make such a loaded, inflammatory statement, at least back it up with some explanation on how you arrived that conclusion.
- viraptor 13y agoThat was put in strong words, but there are some assumptions that are fairly incorrect. About packaging: > Notice that output files of one software package are also inputs (as the filesystem) to the other software packages. Yes, and no. If you're serious about producing the package to distribute it, rather than to install locally, you're going to build it in an isolated environment that has a completely controlled set of dependencies available. For deb packaging that's provided by pbuilder. This is comparable to what happens in nixos at build stage. About deployment: > Our tool needs to: - figure out what kind of package manager we're using on our Linux distribution That's not a dynamic thing. This is solved using plugins and the only operations you have are install (including chosen version/upgrade) and uninstall. There are a lot of system-specific options you probably care about, but they come up because you're most likely running a general purpose system. This means desktop user or developer will want 'apt-get install mysql' to do everything and give you a running server at the end, while people doing automated server deployment will want it to install the binaries and stay far away from the config files or restarting anything. > The problematic part of such system is the fact our tool had to connect to the machine and examine all of the edge cases that machine state could be in If you're running at scale, no, you're not connecting and not figuring stuff out. You'll want to run everything locally. There should be no edge cases either. If something doesn't work, go back to your dev environment and figure out what needs to happen. (or to put it another way, why are there edge cases if all servers are set up from the same description to begin with?) > No dependency hell. Packages stored at unique $PREFIX means two packages can depend on two different openssl versions without any problem. It's just about the dependency graph now. And managing security patches is a bit harder suddenly. You have multiple applications possibly using multiple versions of the same library. > Source and Binary best of two worlds. Only if you can turn the source installation off for the whole system. I don't want package installation to trigger a compiler run on a live system. Ever. > Rollbacks. No state means we can travel through time. Execute --rollback or choose a recent configuration set in GRUB This is not true for any non-trivial system. Can you rollback your database? Maybe, but it may not be able to read the files anymore. Can you roll back your language runtime? Maybe, depends if your code is using new features. Can you rollback a library? Depends what has been built on top of it already. Rolling back the binaries is what existing packages already provide and it's the easiest part of the rollback. So after all this thing about separating yourself from the OS assumptions and various states, etc. we end up with "mkdir -p ${cfg.stateDir}/logs" - why is that embedded into the package at all? What if I'm running on a R/O system and log over network? The article does raise some great points, but then describes a system that either still has the same problems or trades them for something equally bad. Also we still need something to tell all the new servers what the "cfg.stateDir" and other inputs should be. And it will need to know that it's running on nixOS. There are tools that can do that: chef, puppet, salt, ansible, ...
- syongar 13y agoState is the entire value and utility of a computer.
- crashandburn4 13y agook, can you state your argument as to why you think this is the case?
- tomp 13y agothe name "computer" suggests otherwise...
- leoc 13y agoFire is the entire value and utility of, say, a pre-Industrial Revolution forge or bakery, but it doesn't follow that therefore the more fire, or the more uncontrolled the fire, the better: https://en.wikipedia.org/wiki/Great_Fire_of_London https://en.wikipedia.org/wiki/Great_Fire_of_London .
- shaggy 13y agoThe linked article is about package management, not configuration management. Whoever set the title of this post didn't understand the point of the article. From the comments, people seem to confuse and conflate configuration management, job automation and package management. To run a successful infrastructure at any scale you need all three.
- pmahoney 13y agoWhat's the difference between a "package" and a "configuration"? At some level, they are both sets of files, so I can imagine a tool handling both, providing the benefit of a common interface. On the other hand, why might one want separate tools for packages and configuration?
- sp332 13y agoA package is more loosely defined. You can have the same Apache package installed with vastly different configurations. A configuration is a lot more specific.
- jzwinck 13y agoA package is a product, a cooked thing like Nginx. A configuration is ephemeral and site-specific, a dynamic thing like ~/.bashrc. You can claim the two are really the same, but no one else will understand what you are on about (is he saying we should hard-code more stuff?). We don't need exactly two categories here, but that seems the most idiomatic, and has done for decades. Individual users may benefit from versioning their config files, but usually not their programs. Big business might do both, or just as likely screw it up and version code but not configs (I'm looking at you, crontab).
- peterwwillis 13y agoSome people 'cook' their configuration into packages because they aren't allowed to change the configuration on the target. You can also generalize 'automation' to mean configuration management, package management, job management, bug management, change management, deployment, monitoring, continuous deployment/delivery, continuous integration, and many other fields. In practice, everybody screws something up, whether individually or part of a small/big business. The tools are typically not the problem; the humans and how they use them, are.
- kaivi 13y agoTalking about automating apt-get, yum and the like, is there a way to cache frequently downloaded packages on developer machine in the same local network? For instance, I have a bunch of disposable VMs, and I don't want them to download the same gigabytes every time I run init-deploy-test.
- unethical_ban 13y agoapt-mirror creates a full repository, and there's a tool called deb-squid-proxy or some combo of those words that lazily caches packages.
- ema_rocca 13y agoYou might want to give a try to apt-cacher-ng. apt-get install apt-cacher-ng
- cdjk 13y agoI use this - and it's great with my slow internet connection. It's totally transparent, and I only notice it if my vm running it is down. And despite the name, newer versions work with yum too (although I found I had to disable the fastest mirror yum plugin to get reliable caching).
- manish_gill 13y agoAnyone have a good introductory article about these tools (and others like Vagrant etc)? I keep hearing about them, but so far, have been managing a single VPS fine with ssh+git, with supervisord thrown in. Am I missing out by not using these?
- grey-area 13y agoIf you ever manage more than one box, or rebuild your box with very similar configs, Ansible is a great fit and is easy to get started with.
- geerlingguy 13y agoI'm working on Ansible for DevOps[1]. You can download a preview, which has a couple good chapters on setup and initial usage. Even for managing one server, a configuration management tool is a major help; I've started using Ansible to manage my Mac workstation as well[2] (so I can keep my two Macs in pretty much perfect sync). I only assume the reader has basic command-line familiarity, but I try to make the writing approachable for both newer admins and veterans. [1] https://leanpub.com/ansible-for-devops https://leanpub.com/ansible-for-devops [2] https://github.com/geerlingguy/mac-dev-playbook https://github.com/geerlingguy/mac-dev-playbook
- dchuk 13y agoI actually discovered your book yesterday and got very excited about it! What is your target release date for it?
- geerlingguy 13y agoI hope to have the first draft complete by summer. I'm planning on pushing out a new chapter every couple weeks, and already have notes and material for four more chapters... Just need to keep writing. Right now I'm working on a chapter on security, and another on Roles, and will probably publish an updated version next week. I'm also working on getting one or two really good technical editors, as right now I'm basically doing my own editing as I go, one chapter at a time. Once I feel comfortable with where the book is, I'll be publishing on Amazon so people can get a hard copy.
- matlock 13y agoAt least some parts of this post touch on immutable infrastructure, basically just replacing faulty systems and rebuilding them from scratch everytime you need to change it. Relatively easy with AWS and Packer (or other cloud providers) and super powerful. I've written about this a while ago on our blog: http://blog.codeship.io/2013/09/06/the-codeship-workflow-part-4-immutable-infrastructure.html http://blog.codeship.io/2013/09/06/the-codeship-workflow-par...
- vidarh 13y agoIt's easy on your own servers too: Containerise everything or put everything in VMs even when you control the hardware, whether you do it with something like OpenStack or roll your own scripts.
- asuffield 13y agoHow does this system handle shared libraries and security updates to common components? This is not a new idea - the "application directory" dates back to riscos as far as I'm aware. It's been carefully examined many times over the decades, and hasn't been widely adopted because it leads to massive duplication of dependencies, everything in the system has to be changed to be aware of it, and there are less painful ways to solve or avoid the same problems.
- cwp 13y agoIt's probably not a new idea, but it's not the same idea as the application directory as in RiscOS. NixOS doesn't duplicate dependencies. Instead it makes everything read-only. Each dependency is to a specific version of the package, and everything that uses that specific package version uses the same copy. If you want to upgrade to a new version of a package, that implies a new version everything that depends on it as well. (Or at least the application that you want to use the new version.) It still uses more space than a traditional system, but not as much as duplicating everything.
- peterwwillis 13y agoThe more complex the process you use to automate tasks, the more difficult it is to troubleshoot and maintain, and the more impossible it is to inevitably replace parts of it with a new system. https://xkcd.com/1319/ https://xkcd.com/1319/ is not just a comic, it's a truism. I am basically a Perl developer by trade, and have been building and maintaining customized Linux distributions for large clusters of enterprise machines for years. I would still rather use shell scripts to maintain it all than Perl, or Python, or Ruby, or anything else, and would rather use a system of 'stupid' shell scripts than invest more time in another complicated configuration management scheme. Why use shell? It forces you to think simpler, and it greatly encourages you to extend existing tools rather than create your own. Even when you do create your own tools with it, they can be incredibly simple and yet work together to manage any aspect of a system at all. And of course, anyone can maintain it [especially non-developers]. As an example of how incredibly dumb it can be to reinvent the wheel, i've worked for a company that wanted a tool that could automate any task, and that anyone could use. They ended up writing a large, clunky program with a custom configuration format and lots of specific functions for specific tasks. It came to the point where if I needed to get something done I would avoid it and just write expect scripts, because expect was simpler. Could the proprietary program have been made as simple as expect? Sure! But what the hell would be the point of creating and maintaining something that is already done better in an existing ages-old tool? That said, there are certain tasks i'd rather leave to a robust configuration management system (of which there are very few in the open source world [if any] that contain all the functionality you need in a large org). But it would be quite begrudgingly. The amount of times i've ripped out my hair trying to get the thing to do what I wanted it to do while in a time and resource crunch is not something i'd like to revisit.
- rjzzleep 13y agoi both agree and disagree. you can create rather huge complex incomprehensible systems in bash too. also, it's not always better to use existing systems. sometimes rolling something new is indeed better. but i had to work in a place where everyone was a ruby fetishist. don't get me wrong ruby paid my rent for the last couple of years, but i too thought that using it for basic system administration tasks made absolutely no sense. at least what they used it for. it's used in chef/puppet for one reason only. because it's easy to make a dsl in ruby, and it's good for that purpose actually. although i find chef terribly overengineered
- vidarh 13y agoPart of the solutions it to never update "live" machines, but to put everything in VMs, and maintain state outside of the VM images (shared filesystems etc), and build and deploy whole new VM images. Doing updates of any kind to a running system is unnecessarily complex when we have all the tools to treat entire VMs/containers as build artefacts that can be tested as a unit.
- renox 13y agoHis example of replacing the database (stateful) by the network(stateless) for email checking is poor: it makes the implicit supposition that the network is as reliable as the database is.. What happen when one email is lost?
- kinofcain 13y agoWell it's just supposed to illustrate an idea but in this particular example the same thing happens as if you were using the database version and the email with the unique code in it was lost: the user has to request another email be sent. The difference being you don't have to go clean up all the unused tokens out of the db.
- renox 13y agoYes, I realized just after posting this that he was talking about a specific use case (user registration) where the user is the one making the retry in case of failure.. So his example is OK in fact (a variation on the SYN-cookie), I was wrong.
- icebraining 13y agoMore generally, that process depends on both the database and network, so reliability is already the min(database, network) => network. It makes sense to eliminate a component if you can.
- njharman 13y agoAlso the state isn't actually the network, it is somebody's email inbox. Which is just another name for database.
- rbc 13y agoHe left out Cfengine. That's a big gap. It's been around since 1993. He also focused on package management and the provisioning process. I feel like there is more to automation than that. Continuous deployment, process management and distributed scheduling come to mind. As a plus, he does seem to get that just using system images (like Amazon AMI's) can be pretty limited. I think the complexity of automation is more a symptom of the problem space than the tools. It's just a hairy problem. Computer Science seems to largely focus on the single system. Managing "n" systems requires additional scaffolding for mutual authentication and doing file copies between different systems. It also requires the use of directory services (DNS, LDAP, etc…) I like the analogy of comparing the guitar player to a symphony orchestra. When you play the guitar alone, it's easy to improvise, because you don't need to communicate your intent to the players around you. When a symphony does a performance, there is a lot of coordination that needs to be done. Improvisation is much more difficult. That is where Domen is right on target, we can do better. Our symphony needs a better conductor.
- rhys_rhaven 13y agoHe left out Cfengine for the same reason everyone who has used Puppet/Chef/Salt/Ansible leaves it out. Its utterly atrocious. Combining global booleans as a weird sort of declarative flow control to make up for the lack of explicit dependencies between objects that everything else has is horrific. Thousands and thousands of these: (debian_6|debian_7).(role_postgres|role_postgres_exception):: And ordering? Nah, just run it 3 times in a row. :\
- rbc 13y agoIt sounds like you've been using CFEngine 2. CFEngine 3 class (boolean) scope is limited to the bundle it's called from in the bundle sequence. The bundle sequence provides procedural flow control if you need that. Convergent programming does take some getting used to.
- thrill 13y agoCFEngine 3 cleaned up a lot of mess. The promises I write look completely different in it - clean, understandable (even weeks later).
- bdcravens 13y agoI've been playing with Rubber for a Rails app. It's nowhere near as capable as Chef, but for the needs of most Rails apps deploying multiple servers and services to AWS, it's extremely capable. I'd put it somewhere between Chef and Heroku as far as being declarative and being magical.
- mariusmg 13y agoOn Windows i use Boxstarter or a simple powershell script that invokes Chocolatey (must be already installed). I had a look at Puppet/Chef.......wow those really look complicated for something that should really be simple.
- solipsism 13y agoPuppet and Chef do a lot more than install packages, which is literally all that Chocolatey does and most of what Boxstarter does. For example, how would you set a particular value in a configuration file with Chocolatey/Boxstarter? How would you add a system user? How would you make changes to the firewall settings? You can't, that's why Puppet and Chef look more complicated.
- mariusmg 13y agoPowershell. Not to mention that a lot of stuff you would do with Chef/Puppet, on Windows are done directly from the "package" itself (create a default user with limited rights to run under, set up IIS for web apps etc etc). And when something fails....isn't it better to debug your code rather than a black box ?
- noinsight 13y agoActually you're right about Microsoft's solution... PowerShell. But the technology is called Desired State Configuration which is PowerShell based.
- dschiptsov 13y agoWhat problem does it solve besides "I am so clever and just learnt the word 'nondeterministic'?" I would suggest another blog post about monadic (you know, type checked, guaranteed safe) packages (uniques sets of pathnames), statically linked, each file in a unique cryptohashed read-only mounted directory, sorry, volume. Under unique Docker instance, of course, with its own monolithic kernel, cryptohashed and read only. Oh, Docker is a user space crap? No problem, we could run multiple Xens with unique ids.
- telmich 13y agoHas anyone reading this article checked out cdist? I like it very much, you can guess why...
- jes5199 13y agoA thing that this article is hinting at that I think might be more fundamental to making good automation principles: idempotency. Most of unix's standard set of tools (both the /bin programs and the standard C libraries) are written to make changes to state - but automation tools need to assure that you reach a certain state. Take "rm" as a trivial example - when I say `rm foo.txt`, I want the file to be gone. What if the file is already gone? Then it throws an error! You have to either wrap it in a test, which means you introduce a race condition, or use "-f" which disables other, more important, safeguards. An idempotent version of rm - `i_rm foo.txt` or `no_file_called! foo.txt` would would include that race-condition-avoiding logic internally, so you don't have to reinvent it, and bail only if anything funny happened (permission errors, filesystem errors). I does not invoke a solver to try to get around edge cases (e.g., it won't decide to remount the filesystem writeable so that it change an immutable fs...) Puppet attempts to create idempotent actions to use as primitives, but unfortunately they're written in a weird dialect of ruby and tend to rely on a bunch of Puppet internals in poor separation-of-concern ways (disclaimer: I used to be a Puppet developer) and I think that Chef has analogous problems. Ansible seems to be on the right track. It's still using Python scripts to wrap the non-idempotent unix primitives - but at least it's clean, reusable code. Are package managers idempotent the way they're currently written? Yes, basically. But they have a solver, which means that when you say "install this" it might say "of course, to do that, I have to uninstall a bunch of stuff" which is dangerous. So Kožar's proposal is somewhere in the right direction - since it seems like you wouldn't have to ever (?) uninstall things, but it's making some big changes to the unix filesystem to accomplish it, and then it's not clear to me how you know which versions of what libs to link to and stuff like that. There's probably smaller steps we could take today, when automating systems. Is there a "don't do anything I didn't explicitly tell you to!" flag for apt-get ?
- sillysaurus3 13y agoOfftopic, but: what's an example of a situation where using rm -f is bad compared to rm in practice? That is, an example where rm would save you but rm -f would make your life upsetting? On topic: idempotency may be a red herring in this context. Unfortunately filesystems are designed with the assumption that every modification is inherently stateful. (It may be possible to design a different type of filesystem without this assumption, but every filesystem currently operates as a sequence of commits that alter state.) So installing a library or a program is necessarily stateful. What do you do if the program fails to install? Trying again probably won't help: the failure is probably due to some other missing or corrupted state. So indempotency won't help you because there's no situation in which a retry loop would be helpful. That is, if something fails, then whatever operation you were trying to accomplish is probably doomed anyway (if it's automated). I think docker is the right answer. It sidesteps the problem by letting you create containers with guaranteed state. If you perform a sequence of steps, and those steps succeeded once, then they'll always succeed (as long as errors like network connectivity issues are taken into account, but you'd have to do that anyway). EDIT: I disagree with myself. Let's say you write a program to set up a docker container and install a web service. If at some future time some component that the web service relies upon releases an update that changes its API in a way that breaks the web service, then your supercool docker autosetup script will no longer function. The only way around this is to install known versions of everything, but that's a horrible idea because it makes security updates impossible to install. It's a tough problem in general. Everyone agrees that hiring people to set up and manually configure servers isn't a tenable solution. But we haven't really agreed what should replace an intelligent human when configuring a server.
- deleted 13y ago[deleted]
- zobzu 13y agodeterministic builds? pdebuild. mock. this exists since practically forever. as far as the "stateless" thing, this could have been explained in a far simple manner IMO. 1) No library deps: "all system packages are installed in /mystuff/version/ with all their libs, then symlinked to /usr/bin so that we have no dependencies" (that's not new either but it never took off on linux) 2) fewer config deps "only 4 variables can change the state of a config mgmt system's module, those are used to know if a daemon should restart for example" So yeah. it's actually not stateless. And hey, stateless is not necessarily better. It's just less complicated (and potentially less flexible for the config mgmt part). Might be why the author took so long to explain it without being too clear.
- mmcclellan 13y agoThis is an insightful article for devops "teams", That said, a single devop resource can get a hell of a long way in a homogenous Ubuntu LTS environment, apt packaging, Ansible and Github. I know, I know 640k will be enough for anybody, but is anybody's startup really failing because of nginx point releases?
- leftrightupdown 13y agohey, i noticed this on HN so just to share my thoughts. How about incorporating it a bit more simple, just pushing commands with some checks like these guys do? It is a bit more low-level but automation comes only once and should be easier to change. Here is video from their site http://www.youtube.com/watch?v=FBQAhsDeM-s http://www.youtube.com/watch?v=FBQAhsDeM-s
- bqe 13y agoThis is why I'm building Squadron[1]. It has atomic releases, rollback, built-in tests, and all without needing to program in a DSL. It's in private beta right now, but if you're curious, let me know and I'll hook you up. [1]: http://www.gosquadron.com http://www.gosquadron.com
- KaiserPro 13y agoThis isn't stateless, the state has been moved from the package manager/filesystem to a string held in the *INCLUDE. This is nasty.
- sparkie 13y agoThat's not how it works. The Nix build system is effectively a pure function, in that you specify all of the inputs up-front, which includes all dependencies, and then the build system performs the build (which is locally effectful, analagous to say, using the ST monad in Haskell, or Transients in clojure, etc). However, those local effects are discarded when we no longer need the chroot, and the result of the build process, which we want - should be a package that is exactly the same for all inputs. (We're not quite there yet, but approaching this.)
- coherentpony 13y agoShould one really be setting LD_LIBRARY_PATH like that? I thought the preferred way to deal with library search at run time was to rpath it in at compile time.
- jessaustin 13y agoYeah seeing that reminds me of a consulting job I had decades ago. "We're the customer: don't blame us when your software won't run on our servers!" "OK, just set LD_LIBRARY_PATH."
- csense 13y agoI miss DOS, when there was a one-to-one correspondence between applications and filesystem directories. Now Windows programs want to put stuff in C:\Progra~1\APPNAME, C:\Progra~2\APPNAME, C:\Users\Applic~1\APPNAME, C:\Users\Local\Roaming\Proiles\AaghThisPathIsHuge, and of course dump garbage into the Registry and your Windows directory as well. And install themselves on your OS partition without any prompting or chance to change the target. And you HAVE to do the click-through installation wizard because everything's built into an EXE using some proprietary black magic, or downloaded from some server in the cloud using keys that only the official installer has (and good luck re-installing if the company goes out of business and the cloud server shuts down). Whereas in the old days you could TYPE the batch file and enter the commands yourself manually, or copy it and make changes. And God forbid you should move anything manually -- when I copied Steam to a partition that wasn't running out of space, it demanded to revalidate, which I couldn't do because the Yahoo throwaway email I'd registered with had expired. (Fortunately nobody had taken it in the meantime and I was able to re-register it.) I've been using Linux instead for the past years. While generally superior to Windows, its installation procedures have their own set of problems. dpkg -S firefox tells me that web browser shoves stuff in the following places: /etc/apport /etc/firefox /usr/bin /usr/lib/firefox /usr/lib/firefox-addons /usr/share/applications /usr/share/apport /usr/share/apport/package-hooks /usr/share/doc /usr/share/pixmaps /usr/share/man/man1 /usr/share/lintian/overrides I don't mean to pick on this specific application; rather, this is totally typical behavior for many Linux packages. Some of these directories, e.g. /usr/bin, are a real mess because EVERY application dumps its stuff there: $ ls /usr/bin | wc -l 1840 Much of the entire reason package managers have to exist in the first place is to try to get a handle on this complexity. I welcome the NixOS approach, since it's probably as close as we can get to the one-directory-per-application ideal without requiring application changes.
- asperous 13y agoYou should lookup gobolinux
- lewaldman 13y agoWere would be Salt on that bag (http://www.saltstack.com/ http://www.saltstack.com/)?
- eranation 13y agoI'm still failing to understand what solution is there out there that handles web application deployments (especially JVM ones) in an idempotent way. Including pushing the WAR file, upgrading the database across multiple nodes etc. Perhaps there are built-in solutions for Rails / Django / Node.js applications, but I couldn't find a best practice way to do this for JVM deployments. E.g. there is no "package" resource for Puppet that is a "Java Web Application" that you could just ask to be in a certain versions. How do you guys do this for Rails apps? Django apps? is this only an issue with Java web apps?
- druiid 13y agoWell, simply enough, it's kind of a difficult problem to solve. The way most rails apps do it, is deployment with Capistrano or similar. There's also Fabric which can interface pretty well with rails deployment as well. Honestly the methodology I've seen under most deployment situations is far, far from idempotent and is a bit terrible in this respect. There are lots of tools within Rails and Capistrano that hopefully get you to a state approaching 'idempotent' deployments, but they fairly often aren't.
- greatsuccess 13y agoAnother "functional languages make everything better", load of crap.
- greatsuccess 13y agoMaybe we should just ask God to make our servers work. Since we are on the topic of religion anyway.
- Goladus 13y agoThis looks really interesting but I don't see it as a magic bullet for configuration management. There seem to be a lot of advantages on the package management side but configuration management is a lot more than that. Generally the whole point of a configuration file is to allow administrative users to the change the behavior of the application. Treating the configuration file as an "input" is a relatively trivial difference and doesn't really address most of the problems admins face.
- contingencies 13y agoInspired by this post (by a fellow Gentoo user, no less!) I finally published my extended response on the same theme, which has been written over some months: https://news.ycombinator.com/item?id=7384393 https://news.ycombinator.com/item?id=7384393
- gulfie 13y agohttps://www.usenix.org/legacy/publications/library/proceedings/sec96/hollander/ https://www.usenix.org/legacy/publications/library/proceedin... (speaking from second hand knowledge) They don't go into much of the really interesting detail in the paper. The awesome part of all of that was that everything an application required to function was under it's own tree, you never had any question of providence, or if the shared libraries would work right on the given release of the OS you might be using. And it worked from any node on the global network. This problem has been solved, most people didn't get the memo.