11 ms·
Elixir/Erlang Hot Swapping Code (2016)
- omertoast 2y agoi'm so sick of this DevOps bullshit i wonder if there is an alternative language that you can hot swap code and do all the black magic stuff while keeping the reliability and performance like Rust.
- saurik 2y agoI am shocked at the idea of anyone implying that Erlang is "unreliable"... it's entire reason for existence was to step up the game on reliability.
- Me001 2y ago[flagged]
- sgarland 2y agoCan’t wait to hear your take on NodeJS.
- zanderwohl 2y agoI agree. I don't know much about Erlang but what I've heard seems to indicate it's used for high-uptime systems that handle errors well.
- Muromec 2y agoI suspect the causality is reversed. When you have a good designed telecom system, then spmething shaped as erlang happens to be a good tool to create to implement it. The tool than keeps you committed to the design choices you made by being restrictive enough.
- Muromec 2y agoI heard the quote that some 50% of the mobile traffic is handled by erlang. Somehow the other 50% seems to be doing just fine (except the usual shitshow on the inside that sofware is everywhere all the time).
- simoncion 2y ago> I heard the quote that some 50% of the mobile traffic is handled by erlang. Given that you can implement OTP in any language (albeit with varying degrees of difficulty), that's not surprising. The thing to remember is that Erlang was first used in production in like 1986. nearly forty years is more than enough time for the biggest good ideas in Erlang to percolate out into non-BEAM systems.
- LAC-Tech 2y agoAs someone who is now in the rust world and very very sympathetic to the Erlang world... you both probably mean completely different things when you say "reliable". The contexts are just world apart.
- dcsommer 2y agoGP seems to be implying hot swapping, not Erlang, is unreliable. To which, from my experience using it in Erlang, I heartily agree is fraught. Inconsistent state across nodes is much harder to reason about. When you _must_ ensure consistency, hot swapping is reckless, especially as org size and product complexity increases. Leave hot loading to local/development environments, not production deploys. Loading configs on the fly can also have some of this risk, but it is much easier to reason about typically.
- sergiotapia 2y agowe had this wonderful thing in PHP where you would just save a .php file and bada bing it was LIVE. what happened? :D
- thanksgiving 2y agoThey took away our access after one too many outages :yay:
- Muromec 2y agoSounds like what happened with hot upgrade privileges in some erlang shops too.
- thanksgiving 2y agoI've worked at a smaller project where it worked just fine. The key I think is to have the project be small enough to be able to fit in my head. That and have an identical test server. I used to make changes locally, test it locally, then make the same change on test, have someone else look at it, get a lgtm, and do the same thing on the production machine. It sounds like a lot of steps but it is pretty straightforward. Sadly, it probably doesn't work with bigger teams or more complicated projects.
- Muromec 2y agoA lot of things work on a small project where everybody knows what they are doing and can keep all the dependencies and data structure in their heads. Logging to the production server and editing PHP file right where apache is looking at it and fixing stuff with sql commands in the production database to address customer complaints. The big question is what everyone else should be doing that survives the touch with unevenly distributed amounts of technical expertise and amount fucks given about result.
- cardanome 2y agoWe still have that and it is awesome. PHP is better than ever. In serious emergencies I even sometimes end up quickly SSH-en to a prod server and changing the file directly. Which is kind of horrifying but hey customer is happy it got fixed immediately and I get to relax and take my time to write a proper fix. Beats sweating and watching the pipeline build and asking around for people to approve my merge request.
- Muromec 2y agoThe black magic comes at the cost of not having one streamlined procedure to release stuff. And to make the black magic work you have to engage with it. Most of the time people don't even bother write to a proper import.meta.hot.accept thingy in javascript. Developers simply hate chores, which is evident by not willing to write proper unit tests (despite knowing that tests work) or writing just enough to let the coverage cop pass ship the build. A dedicated small team running something like whatsup? Sure, look into the arcane and let it look back at you (although high insight makes one more susceptible to madness you know). But most of the time you will do better job with PHP in a stupid restartable box behind seven load balancing proxies.
- yetihehe 2y ago> The black magic comes at the cost of not having one streamlined procedure to release stuff. You can also have a streamlined procedure to release stuff. Most changes in my erlang based system consist of "push to staging branch, click to deploy and test, pull to master, click deploy button". Can't be simpler than that. Most changes in such systems are also pretty simple. When you need to add something big, typically not many things are dependent on that, so deploy is also pretty simple. > But most of the time you will do better job with PHP in a stupid restartable box behind seven load balancing proxies. Yeah, we talk here about more complicated things here. If you have something simple, you don't need to use erlang, `python -m http.server` will be even simpler than your php in stupid restartable box, because you don't need a special box, just one small command.
- Muromec 2y agoDo you do 100% of deployments using hot reload? If yes, maybe you should share the recipy with everybody else, since consensus seems to recommend the opposite. At the very least you will have a different procedure to upgrade the erlang itself, right? >If you have something simple, you don't need to use erlang I think on a spectrum of difficult things there is an area between hosting static file on rpi at home and running massivele distributed system full of long running stateful processes.
- lamuswawir 2y agoErlang is built for reliability. They're chasing nine nines. Everything about the BEAM is built to emphasize that, the design choices, the documentation, the recommended practices. Erlang is not very fast, but that's not what it was built for.
- Muromec 2y agoIs it really beam or just otp? Sure, beam gives you processes, network-transparent send, immutable structures and linking-monitoring thingy on top, but is what makes it good to shoot for nines? I suspect the aura of mistycism around yet another jit vm is not that warranted
- jerf 2y agoIt is reasonable to conceive of Erlang as encompassing OTP. Perhaps somewhere in the world there is some developer out there hot on Erlang but just hates OTP and doesn't use it, but they must be fairly frustrated at how hard it is to keep OTP out of their code base if they ever need any libraries. Restarting is arguably the definitive thing that makes Erlang stack the 9s out past what most languages and runtimes can achieve... the thing is, it's more complicated to use in practice than a web page like this makes it look, and it's beyond what most products need. Few applications need the fifth or sixth or seventh nine, and it gets to the point that you can't have it anyhow because your Erlang cluster, no matter how well distributed, itself probably doesn't have 99.99999 availability, and your users probably don't have 99.99999 availability on their own network connection. It's not impossibly complicated, but it is the sort of thing where you if you want to use the feature you need to have it sort of constantly in mind as you write the rest of your system, and it's a lot easier even in Erlang to just design the system to take entire nodes down and bring them back up, if not the entire cluster down, rather than fuss with hot reloads. I wish Erlang advocates would be more upfront about pitching this as an interesting niche feature, but not really a reason to consider Erlang. Unless you absolutely need it, in which case it can indeed be the thing that puts it on the short list of choices... but as evidenced by the vast, vast majority of software and systems not being on Erlang and managing to get along, there aren't really that many things that need it.
- gf000 2y agoIt's not well known, but the JVM has very good hot reload support, and is a very reliable and performant platform.
- amelius 2y agoDoes this hot swapping also work for closures?
- Muromec 2y agoErlang doesn't have closures, because erlang doesn't have variables. The compiler simply desugars it to partially applied function referenced by it's name (yes, those inline functions in fact have names). If you have something_function, then first inline function used in it will be -something_function/1-fun-0- with zero being the index and captured variable being another argument. Now if you will change the host function to have more inlines before it, the indexing will drift. So I would expect the body of inline function will still be resolved from the old version of the module, but I didn't actually try. Source: I did run erlc -S at least once. Add: now thinking of it, will the call to a local function from the old version of the module ever escape into the new one without first returning back to gen_server and letting it call the new version? Another comment says that calls withing the module never do, so the assumption was correct.
- bitwalker 2y agoErlang absolutely has closures, you are mistaken. What you are referring to are "function captures", which bind a function reference as a value, and there is no environment to close over with those. However, you can define closures which as you'd expect, can close over bindings in the environment in which the closure is defined. The interaction between hot reloads and function captures in general is a bit subtle, particularly when it comes to how a function is captured. A fully qualified function capture is reloaded normally, but a capture using just a local name refers to the version of the module at the time it was captured, but is force upgraded after two consecutive hot upgrades, as only two versions of a module are allowed to exist at the same time. For this reason, you have to be careful about how you capture functions, depending on the semantics you want.
- toast0 2y ago> but is force upgraded after two consecutive hot upgrades, as only two versions of a module are allowed to exist at the same time. Force upgraded is maybe misleading. When a module is loaded for the 3rd time, any processes that still have the first version in their stack are killed. That may result in a supervisor restarting them with new code, if they're supervised.
- slt2021 2y agohot reload of code is nothing new nowadays, but people use it only locally during development for REPL like development style. in actual production, people prefer to operate at the container level + traffic management, and dont touch anything deeper than the container
- foota 2y agoAmusingly, this reminds me sort of about the story of a person who joins a new company only to discover that their programming framework is intricately linked to their version control system.
- myfavoritedog 2y ago[dead]
- diath 2y ago> in actual production, people prefer to operate at the container level + traffic management, and dont touch anything deeper than the container How do you think video games like World of Warcraft or Path of Exile deploy restartless hotfixes to millions of concurrent players without killing instances? I don't think it's a matter of "prefer to", it's a matter of "can we completely disrupt the service for users and potentially lose some of the state"? Even if that disruption lasts a mere millisecond, in some context it's not acceptable.
- anonymousDan 2y agoI'm a distributed setup I imagine there could be cases where you want to atomically hot upgrade multiple VMs at the same time. Is this common in practice and if so are there recommended patterns/techniques for doing it?
- AlphaWeaver 2y agoErlang does have a mechanism that allows a module to control when it moves from the "old version" to the "new version" of its own code. Calls to the module with the fully qualified name (e.g. `module:function()`) will invoke the "new code" once it's loaded, but calls within that module using only function names (just `function()`) will continue to invoke the "old code". If the portion of the app you were hot upgrading was an OTP process like a GenServer, you could theoretically wait for some sort of atomic coordination mechanism to make that fully qualified function call after the new code has loaded, at least in theory. We use hot code reloading at my work, but haven't had a reason to atomically sync the reload. Most of the time it's a tmux session with `synchronize-panes` and that suffices. If your application can handle upgrades within a module smoothly, it's rare to have a need for some sort of cluster-level coordination of a code change, at least one that's atomic.
- Muromec 2y agoThere can't be anything atomic in a distributed system. You can't even atomically hot upgrade it on a single VM anyway -- you instead load the new version of the module and let dispatcher know to route new calls into it, the same as you would do with a load balancer and a bunch of load bearing docker hosts, just inside your app.
- knome 2y agoerlang has a code_change function in the otp that allows the gen_server to update its current state and start using new code. No connections need be broken with clients, no long running processes need be stopped. Just updated in place. It's not just a routing change. https://www.erlang.org/docs/24/man/gen_server https://www.erlang.org/docs/24/man/gen_server
- behnamoh 2y agoLisp has had this features since day 1. But Lisp-like langs like Clojure, Racket, etc. don't have it. This is one of the fundamental features of Common Lisp and I don't know why most other Lisp-wanna-be's don't implement it.
- lamuswawir 2y agoCame here to say this. In Lisp, you can just compile a function, or load a file and it just works. It's not even sold as a hot feature, not the way Erlang sells it. It's just a feature. I manage a few websites written in Lisp, and updating them is as simple as push code, recompile and it works.
- davidw 2y agoBut what if the system is running and the new function takes different arguments or something? What if there is data loaded in the system, what happens to it? Simply loading new code is easy, ensuring the whole system works seems to require a bit more effort.
- fiddlerwoaroof 2y agoCommon Lisp has a bunch of features designed to enable migrating the system. e.g. update-instance-for-redefined-class ( https://www.lispworks.com/documentation/HyperSpec/Body/f_upda_1.htm https://www.lispworks.com/documentation/HyperSpec/Body/f_upd... ) lets you write code to update instance data between class versions when a class definition is reloaded. It turns out, though, that making hot-code reloading work well is mainly a question of how you design your system: designing for hot code reloading isn't all that hard for 90% of cases once you figure out the relevant techniques.
- deleted 2y ago[deleted]
- leprechaun1066 2y agoWe do this in q/kdb+ systems often for patches. An important thing about these languages is that this kind of workflow is part of the core for solving problems. So when you are building a system one of the aspects of its design will always allow for this update method. Then when you push a patch you both know the impact of the change (because you've tested the exact same steps in a dev/QA/UAT/Beta environment) and the work required to do it safely. Major releases do go through a full shutdown and release cycle though.
- Volundr 2y agoIt's worth noting that distillery is deprecated in favor of mix releases, which don't support relups out of the box, and specifically warn against them due to the complexity involved in writing code to support them correctly. It's a cool feature that's no doubt amazing for applications that need it, but it brings a fair amount of complexity vs other deployment strategies.
- superdisk 2y agoYeah, note that this article is from 2016. I distinctly remember during that time that these hot-swap deployments were all the rage in the Elixir community, and then fell out of fashion with time.
- thibaut_barrere 2y agoGood point. Someone shared this in case someone wonders: https://elixirforum.com/t/how-to-tweak-mix-release-to-work-with-hot-code-reloading/65188 https://elixirforum.com/t/how-to-tweak-mix-release-to-work-w... > I’ve spent some time understanding how to do hot code reloading with releases built using mix release, and here I’d like to detail the steps needed, in hopes that it will help someone.
- alberth 2y ago(2016)
- dang 2y agoWhere do you see that? I couldn't find it.
- gnabgib 2y agoIt's in the URL :D But yeah, the page doesn't make it clear (and some of the embedded JS has a 2020 date suggesting it's received updates). In the RSS feed too: Wed, 07 Dec 2016 https://kennyballou.com/index.xml https://kennyballou.com/index.xml
- dang 2y agoHidden in plain view! Ok, let's put 2016 above, on the assumption that the edits since then haven't been too major.
- hauxir 2y agoAt kosmi.io we use elixir hot swapping for every small patch/bugfix on the backend. This allows us to deploy updates multiple times a day with 0 disruption. Allows the clients to remain connected and be none the wiser that there was an update at all. For larger updates we just do hard restarts when in-memory data structures or supervision tree are changed.
- deathtrader666 2y agoWould love to know more how you go about it.
- hauxir 2y agoIt's a little hacky but I'll try to explain: * The server runs in a docker container which has an ssh server installed and running in the background. The reason for SSH is simply because that's what edeliver/distillery uses. * The CI(local github runner) runs in a docker container as well which handles building and deploying the updated releases when merged on master. * We use edeliver to deploy the hot upgrades/releases from the CI container to the server container. This happens automatically unless stopped which we do for larger merges where a restart is needed. * The whole deployment process is done in a bash script which uses the git hash for versioning, edeliver for deploying and in the end it runs the database migrations. I'm not going to say it's perfect but it's allowed us to move pretty damn fast.
- deleted 2y ago[deleted]
- GCUMstlyHarmls 2y agoThis is a talk about a large scale, resilient elixir/erlang deployment in healthcare. Specifically they talk about running with no down time using hot code reloading here: https://youtu.be/pQ0CvjAJXz4?t=2667 https://youtu.be/pQ0CvjAJXz4?t=2667 but the whole talk is quite interesting regarding availability. Warning: the video is quite quiet.
- benzible 2y ago"hot deploys on fly.io to a planet-wide cluster, in 3 seconds.": https://x.com/chris_mccord/status/1785678249424461897 https://x.com/chris_mccord/status/1785678249424461897
- jongjong 2y agoForcing all clients to reload their code at the same time sounds like a bad idea. Allowing different clients to run different incompatible versions of the code at the same time also sounds like a bad idea. APIs are like database engines; they should rarely change. Making it easy to change them is an anti-pattern. Engineers don't build bridges with replaceable pillars or skyscrapers with replaceable foundations. When aerospace engineers tried building a plane with replaceable engines, we got Boeing 737 Max...
- tzmudzin 2y agoEngine replacement happens on airplanes fairly frequently. You don't want to scrap an airplane because of a single damaged turbine blade, or even keep it on the ground for longer. https://jalopnik.com/how-airlines-decide-to-replace-jet-engine-boeing-airbus-1850275010#:~:text=The%20plane's%20jet%20engines%20will,happens%20after%2012%2C000%20flight%20cycles https://jalopnik.com/how-airlines-decide-to-replace-jet-engi....
- jongjong 2y agoYes but the new parts meet the specs of the original design. The design itself isn't flexible. You can't make the engines significantly bigger without significantly revising the blueprint as a whole. That was the Boeing Max lesson. Just changing the software was not enough.
- p_l 2y ago737 MAX had nothing to do with replaceable engines, but with trying to run an ancient airframe with new engines but without necessary upgrades to support the new engines because of costs.
- jongjong 2y agoReplaceable at the design level. OMG. Why do I have to explain everything? Clearly I'm talking about blueprints here. Code is a blueprint since you can launch multiple processes/instances running the same code.
- apex_sloth 2y agoI used to work for a company that wanted zero downtime through Erlang's hot code reload feature. While it absolutely works, it requires immense effort and extra code to handle state upgrades and downgrades.
- modernerd 2y agoLive updating a drone running Erlang in 10ms while it was flying with no application restart and no loss of state impressed me when I saw it in 2021: https://www.youtube.com/watch?v=XQS9SECCp1I https://www.youtube.com/watch?v=XQS9SECCp1I But I almost never hear Erlang/Elixir/Gleam folks talk about this benefit of the Erlang VM now, even though it seems fairly unique and interesting. Has the community moved away from it? Is it just not that useful?
- cess11 2y agoA lot of the GenServer-information floating around explains code_change/3, no? That's commonly what you want, a way to handle state propagation when process code is updating in a running system. Most people are probably running some web services or something and might as well shift machines in and out of a cluster or can wait for old processes to disband on their own, because the new code is backwards compatible with the one in already running processes, and so on. It can also be relatively hard to do without causing damage to the system. Those who need and can manage it probably don't need it marketed.
- deleted 2y ago[deleted]
- cess11 2y agoSomeone put a reply and then deleted it while I wrote a response, and it irks me that it might have been a waste so here's the gist of it: "Is it just that people are more comfortable with blue-green deploys, or are blue-green deploys actually better?" It depends. If you can do a blue-green shift where you gradually add 'fresh' servers/VM:s/processes and drain the old, that's likely to be most convenient and robust in many organisations. On the other hand, if you rely on long running processes in a way where changing their PID:s break the system, then you pretty much need to update them with this kind of hot patching. "Does Erlang offer any features to minimize damage here?" The BEAM allows a lot of things in this area, on pretty much every level of abstraction. If you know what you're doing and you've designed your system to fit well into the provided mechanisms the platform provides a lot of support for hot patching without sacrificing robustness and uptime. But it's like another layer of possible bugs and risks, it's not just your usual network and application logic that might cause a failure, your handling of updates might itself be a source of catastrophe. In practice you need to think long and hard about how to deploy, and test thoroughly under very production like conditions. It helps that you can know for sure what production looks like at any given time, the BEAM VM can tell you exactly what processes it runs, what the application and supervisor trees look like, hardware resource consumption and so on. You can use this information to stage fairly realistic tests with regards to load and whatnot, so if your update for example has an effect on performance and unexpected bottlenecks show up you might catch it before it reaches your users. And as anyone can tell you who has updated a profitable, non-trivial production system directly, like a lot of PHP devs of ye olden times, it takes a rather strong stomach even when it works out fine. When it doesn't, you get scars that might never fade.
- melvinroest 2y agoIs this like a similar feature in Smalltalk/Pharo and Lisp?
- igouy 2y agoYes, the basics are there in Smalltalk and there's more support built into Erlang. Also: "Live program changes in the Dart VM" https://github.com/dart-lang/sdk/blob/main/docs/Hot-reload.md https://github.com/dart-lang/sdk/blob/main/docs/Hot-reload.m... "Live reloading for your ESP32" https://github.com/toitlang/jaguar https://github.com/toitlang/jaguar
- gregors 2y agoThe Big Elixir 2018 - Desmond Bowe - Hot Upgrade Are Not Scary https://www.youtube.com/watch?v=IeUF48vSxwI https://www.youtube.com/watch?v=IeUF48vSxwI
- epiccoleman 2y agoI wonder if this kind of thing could be used to make the Elixir REPL a bit more LISPy. I like iex a good deal, but I often find myself wishing I could just easily eval some code or expression in the editor and have it make its way into the REPL context. (yes, I know you can `r` on a module, but that's pretty clunky compared to something like CIDER).
- dszoboszlay 2y agoHot code upgrades on the BEAM are awesome, but they're not a piece of cake. If you're also interested in the challenges of making them production safe, I gave a talk about this topic on CodeBEAM Sto earlier this year: https://youtu.be/epORYuUKvZ0?si=gkVBgrX2VpBFQAk5 https://youtu.be/epORYuUKvZ0?si=gkVBgrX2VpBFQAk5 OP talks in the summary about the importance of understanding the process. It's very much true, but you need to understand not only the process your tooling provides, but also what's going on in the background and what hasn't been taken care for you by your tools. I'm afraid these things are rarely understood about hot upgrades, even by experienced Erlang engineers.
- robocat 2y agoGreat discussion 23 days ago on hot code loading: https://news.ycombinator.com/item?id=42187761 https://news.ycombinator.com/item?id=42187761