3 ms·
The big thing about immutable infrastructure is that it is reproducible. I've seen both worlds and I do appreciate the simplicity and quickness of the upgrade s
by Sebb767 2y ago
The big thing about immutable infrastructure is that it is reproducible. I've seen both worlds and I do appreciate the simplicity and quickness of the upgrade solution presented in the post. The problem with this manual approach is that it is quite easy to end up with a few undocumented fixes/upgrades/changes to your pet server and suddenly upgrading or even just rebooting the servers/app becomes something scary.
Now, for immutable infrastructure, you have a whole different set of problems. All your changes are nicely logged in git, but to deploy you need to rebuild containers and roll them out over a cluster. To do this smoothly, the cluster also needs to have some kind of high availability setup, making everything quite complex and, in the end, you wasted minutes to hours of compute for something that a pet setup can do in a few seconds. But you can be sure that a server going down or a reboot are completely safe operations.
What works for you really depends on your situation (team size, importance of the app, etc.), but both approaches do have their uses and reducing the immutable infra approach to "people run k8s because it's hip" misses the point.
- swiftcoder 2y ago> undocumented fixes/upgrades/changes to your pet server and suddenly upgrading or even just rebooting the servers/app becomes something scary. You can mostly prevent this by mandating that fresh nodes come up regularly. Have a management process that keeps a rolling window of ~5% of your fleet in connection-drain, and replaces the nodes as soon as they hit low-digits of connections. Whole fleet is replaced every ~3 weeks, you learn about any deployment/startup failures within one day of new code landing in trunk, minimal disruption to client connections.
- turtlebits 2y agoThis doesn't prevent anything, it just schedules possible breakages because your infra isn't 100% immutable. IME, this doesn't work because companies won't implement/will deprioritize any infra changes that impact the development cycle.
- swiftcoder 2y agoIt's not so very different to dropping your PR into any other automated-CI/CD-all-the-way-to-prod pipeline. Albeit maybe a little easier to justify to management that you are dropping everything to fix the breakage when your CI/CD pipeline stops.
- packetlost 2y ago> Now, for immutable infrastructure, you have a whole different set of problems. All your changes are nicely logged in git, but to deploy you need to rebuild containers and roll them out over a cluster. The real issue is that it effectively forces externalizing nearly all state. On the surface, this seems like it's just a good thing, but if you think about the limitations and complexity it creates, it starts seeming less unquestionably good. Sometimes that complexity is warranted, but very frequently it is not. That being said, I think modifying code is a running system without pretty strict procedures/control around it is... dangerous. I've seen hotfixex get dropped/forgotten because it only existed on running system and not in source control more than a couple of times.
- toast0 2y ago> The big thing about immutable infrastructure is that it is reproducible. I've seen both worlds and I do appreciate the simplicity and quickness of the upgrade solution presented in the post. The problem with this manual approach is that it is quite easy to end up with a few undocumented fixes/upgrades/changes to your pet server and suddenly upgrading or even just rebooting the servers/app becomes something scary. The point is not really automation vs manual. Hot loading is amenable to automation too. The point is really that when you replace immutable servers with state with another set, there's a lengthy process to migrate the state. If you can mutate the servers, you save a lot of wall clock time, a lot of server cpu time, and a lot of client cpu time. I deal with this issue at my current job. I used to work in Erlang and it took a couple minutes to push most changes to production. Once I was ready to move to production, it was less than 30 minutes to prepare, push, load, verify and move on with my life. I could push follow ups right away, or wrap up several issues, one at a time, in a single day. Coming from PHP was pretty similar, with caveats about careful replacement of files (to avoid serving half a PHP file) and PHP caching. Now I work with Rust, terraform, and GCP; it takes about 12 minutes for CI to build production builds, it takes terraform at least 15 minutes to build a new production version deployment, and several more minutes for it to actually finish coming up, only then can I start to move traffic, and the traffic takes a long time to fully move, so I have to come back the next day to tear down the old version. I won't typically push a follow up right away, because then I've got three versions running. I can't push multiple times a day. If I'm working many small issues, everything has to be batched into one release, or I'll be spending way too much of my time doing deploys, and the deployment process will be holding back progress.
- fiddlerwoaroof 2y agoThe funny thing here is BEAM is “immutable infrastructure as a programming language environment” which, to me, is strictly superior to the current disjunction between “infrastructure configuration” and “application code”. Erlang defaults to pure code and every actor is like a little microservice with good tooling for coordination. There are mutable aspects like a distributed database, but nothing all that different from the mutable state that exists in every “immutable infrastructure” deployment I’ve seen.