5 ms·
In practice, hot code reloading and mnesia are pretty cool concepts but very far from trivial. The vm and OTP provide mechanisms to work with such advanced conc
by rkachowski 3y ago
In practice, hot code reloading and mnesia are pretty cool concepts but very far from trivial. The vm and OTP provide mechanisms to work with such advanced concepts, but it still requires an intense amount of discipline and process to execute effectively.
Consider hot code reloading to be like changing the tyres on a moving car, the pieces and tools are there to make it possible, but in almost every situation it's more sensible to stop and change. Backwards compatibility for all API changes, code correctness and deployment methodology need to be fixed before you can start a hot rollout. That's not even to consider a flexible application architecture that will be compatible with such changes in functionality (OTP does indeed ensure this is compatible on the basic process level, but there is no magic bullet around high level architecture).
For given hard requirements, e.g. upgrading the application without dropping TCP connections - this can be a killer feature that removes considerable development effort, but usually it's a very cool thing that you don't need.
That being said, hot code reloading on the individual module level I have found to be a useful and dirty trick to change the behaviour of running applications (add log statements / telemetry where they weren't before to diagnose production issues). This still has all the caveats and danger of the above, but is more limited in scope.
- dikei 3y agoExactly, it's similar to Linux live kernel patching. It's certainly desirable, but with the support of modern orchestration platform, most of the time it's simpler to migrate services off the nodes, update, reboot then redeploy.
- toast0 3y agoHaving worked with hot loading for a long time, and now working without, even if you bring up new nodes, migrate services, destroy old nodes, it just takes so much longer, and you have to spend so much more time dealing with the multiple versions at once problem. Especially if you have state in your service. Shutting down servers with active clients is disruptive, and moving clients immediately is difficult. Simple HTTP request/response is easier, but a lot depends on your load balancing.
- tetha 3y ago> Consider hot code reloading to be like changing the tyres on a moving car, the pieces and tools are there to make it possible, but in almost every situation it's more sensible to stop and change. Backwards compatibility for all API changes, code correctness and deployment methodology need to be fixed before you can start a hot rollout. That's not even to consider a flexible application architecture that will be compatible with such changes in functionality (OTP does indeed ensure this is compatible on the basic process level, but there is no magic bullet around high level architecture). This was an interesting discussion I had with a few of our B2B customers at a customer event. Even with redundancy and everything, something like a major postgres upgrade is a tricky thing. You either have to do some amount of work of setting up several clusters, setup logical replication, decide a cutover point (most of them things pg_easy_replicate (if I remember the name right) does), while juggling the backup solution on the side and such. It's overall understandable, but a decent number of steps to do right, and if you don't you might even lose transactions to the upgrade. And losing transactions is way worse than being down. The alternative though is somewhat brute: After testing everything, shut the app down, shut the cluster down, upgrade binaries, run pg_upgrade. Restart cluster, check cluster, start application. This requires downtime, but as a procedure, it is very straightforward, very testable and overall predictable. Running a readonly version of the application sadly wouldn't be possible because of choices in the application. Interestingly, many of our larger customers obviously preferred if vendors don't go down, naturally. But for a complex migration or update, most of them have been burnt by horror stories of upgrades going entirely sideways. If an outage is of predictable time, well-planned and communicated in advance, they can prepare for it and deal with it. One just shrugged and said "So what, then we just tell our customers we don't provide that service for half a day. No big deal with a month in advance or so" This gave me a very valuable perspective of how much energy to put into avoiding downtimes for rare procedures and interactions. Talk with your customers, and sometimes the correct answer is indeed "don't".
- darkmarmot 3y agoWe use hot code reloading because taking our system down for any length of time could kill people (real time medical software).
- darkmarmot 3y agoWe run a distributed cluster of 32 Elixir nodes across 4 data centers. Typically, we do rolling deployments -- but when we need a fast fix for something, pushing patches against the live system is real life saver!