4 ms·
When I worked at Discord, we used BEAM hot code loading pretty extensively, built a bunch of tooling around it to apply and track hot-patches to nodes (which in
by jhgg 2y ago
When I worked at Discord, we used BEAM hot code loading pretty extensively, built a bunch of tooling around it to apply and track hot-patches to nodes (which in turn could update the code on >100M processes in the system.) It allowed us to deploy hot-fixes in minutes (full tilt deploy could complete in a matter of seconds) to our stateful real-time system, rather than the usual ~hour long deploy cycle. We generally only used it for "emergency" updates though.
The tooling would let us patch multiple modules at a time, which basically wrapped `:rpc.call/4` and `Code.eval_string/1` to propagate the update across the cluster, which is to say, the hot-patch was entirely deployed over erlang's built-in distribution.
- davisp 2y agoThis matches my experience. I spent a decade operating Erlang clusters and using hot code upgrades is a superpower for debugging a whole class of hard to track bugs. Although, without the tracking for cluster state it can be its own footgun when a hotpatch gets unpatched during a code deploy. As for relups, I once tried starting a project to make them easier but eventually decided that the number of bazookas pointed at each and every toe made them basically a non-starter for anything that isn’t trivial. And if its trivial it was already covered by the nl (network load, send a local module to all nodes in the cluster and hot load it) style tooling.
- SexyKitty17 2y ago[flagged]
- SexxxyKitty17 2y ago[flagged]
- scotty79 2y ago> Although, without the tracking for cluster state it can be its own footgun when a hotpatch gets unpatched during a code deploy. This and everything else said sounds so much like PHP+FTP workflow. It's so good.
- stouset 2y agoCan someone explain how this is not genuinely terrifying from a security perspective?
- aunderscored 2y agoIt's the same amount of terrifying as a regular deploy, you need to ensure that you limit access as needed
- nelsonic 2y agoWhere is the security problem? All code commits and builds can still be signed. All of this is just a more efficient way of deploying changes without dropping existing connections. Are you suggesting that hot code replacement is somehow a attack vector? Ericsson has been using this method for decades on critical infrastructure to patch switches without dropping live calls/connections it works. No need to fear Erlang/BEAM.
- stouset 2y agoMy interpretation of the GP was that a code change in one node can be automagically propagated out to a cluster of participating Erlang nodes. As a security person, this seems inherently dangerous. I asked why it is safe, because I presumed I’m missing something due to the lack of ever hearing about exploitation in the wild.
- badpenny 2y agoWhy is it any more dangerous than a conventional update, which also needs to be propagated?
- stouset 2y agoA conventional update takes place out of band. If someone were to exploit a running Erlang process, the description of this feature sounds to me like they would have access to code paths that allow pushing new code to other Erlang processes on cooperating nodes.