4 ms·
I've used Elixir/Phoenix on K8s and there's some truth to this. For example - kubernetes' "if the process dies, kill the whole container and restart it" model d
by hanrelan 9y ago
I've used Elixir/Phoenix on K8s and there's some truth to this. For example - kubernetes' "if the process dies, kill the whole container and restart it" model doesn't work very well with Elixir (a process within the VM can die and K8s will have no idea). To some degree, K8s is replicating some of the BEAM functionality.
But there's still a lot from Kubernetes that Elixir can take advantage of. For example - service discovery, autoscaling, clustering, secret management - all made easier with a K8s deployment and Elixir.
I wrote a few blog posts about it a while ago if you're interested: https://medium.com/polyscribe/a-complete-guide-to-deploying-elixir-phoenix-applications-on-kubernetes-part-1-setting-up-d88b35b64dcd https://medium.com/polyscribe/a-complete-guide-to-deploying-...
- sb8244 9y agoHmm I'm not quite sure I follow the process examine. If your started process crashes (let's say supervisor fails to start due to too many failures), k8s would restart the process (your app) to try again. If a managed process in elixir dies, you could make that bring down your whole app, but generally you would want to handle it yourself. This is similar to an app server in Ruby. A single worker might die, but it's still being managed by a central coordinator. If the coordinator dies, everything is down and the app restarts
- dguaraglia 9y agoNot trying to be flippant, but by extending your analogy: if the Kubernetes supervisor dies (which is conceivable, no technology is infallible), then you are in the same scenario as something bringing your whole Elixir VM down. You can always run your Elixir VM using something like supervisord if you need another watchdog. BTW, it's very hard to bring BEAM down. Most exceptions will bring kill the faulty process, but leave the supervisor tree untouched. Only "unmanaged" code (for example, calling a C library using a NIF) might crash the whole VM. Then again, no technology is infallible.
- sb8244 9y agoTrue true, maybe there's just a single part that I was hung up on in parent comment > (a process within the VM can die and K8s will have no idea) I suppose that I see nothing wrong with this and would say that it's not weird or unusual to be the case. They're separate things, although they both achieve visually similar goals (keep things running as much as possible). My point is less about K8s vs Elixir but rather a statement that this seems acceptable and shouldn't be a reason to rule out something like running Elixir on K8s.
- greenleafjacob 9y ago> You can always run your Elixir VM using something like supervisord if you need another watchdog. Sounds like heart: http://erlang.org/doc/man/heart.html http://erlang.org/doc/man/heart.html
- noemotion 9y ago>BTW, it's very hard to bring BEAM down. Most exceptions will bring kill the faulty process, but leave the supervisor tree untouched. Only "unmanaged" code (for example, calling a C library using a NIF) might crash the whole VM. Not completely true. Try to allocate more memory than available and then see what happens.
- dguaraglia 9y agoI'm actually curious. Do you have some first-hand experience with that kind of issue? I've only played around with Elixir and my only experience deploying BEAM-based applications on production has been RabbitMQ and CouchDB and both of those were very solid. Also, isn't ENOMEM a standard error that you can handle using the standard error handling techniques?
- noemotion 9y agoI have and it wasn't good. It is impossible to handle out of memory exceptions inside the BEAM. At best, you can hope for a crash of the VM. At worse, it just hangs. And one will most likely not find out about this until they hit this scenario.
- hanrelan 9y agoThe scenario I saw with this was the following: BEAM VM comes up and tries to connect to the database. The database is temporarily unreachable from that machine, but only the Ecto process dies and BEAM VM doesn't die. Since the last process (BEAM) in the Dockerfile is running, Kubernetes thinks it's healthy and will try sending traffic to it. Compare this to something like node, where the node process will die if it can't connect to the database, and Kubernetes won't try to direct traffic to it and will restart it for you. You can do this with Elixir supervisors as well - that's why I said Kubernetes replicates some of Elixir's functionality. It's definitely possible to work around all this which is why we ended up using Kubernetes for our Elixir deployment (and I'd recommend it). It's just these little things that make it clear that Kubernetes was designed with something like node in mind and not BEAM.
- jclulow 9y agoThe behaviour you are attributing to Node is really just the behaviour of poorly constructed software. There are at least some Node-based software products which correctly come up and wait for dependencies to become available -- and gracefully handle their subsequent transient unavailability.
- atopuzov 9y agoThere are 2 types of probes: readiness and liveness [1]. Define your readiness probe so it passes when you are ready to take the traffic (eg. connected to the db). [1] https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-probes/ https://kubernetes.io/docs/tasks/configure-pod-container/con...
- josevalim 9y agoUnless Kubernetes is killing the container without respecting the shutdown time of the underlying processes, this should not be problem. Once Kubernetes sends an exit signal to the VM, the VM will start the shutdown of its processes, and everything should terminate gracefully. It is OK that Kubernetes doesn't know about Elixir processes terminating, because it is not its job. It is the job of the Elixir software to express its start-up and shutdown guarantees. If you don't want the VM to start when you don't have a database connection, then you should start a process right after your connection pool starts and before your endpoint runs that attempts to issue a query and act accordingly.
- rkangel 9y agoThat's a very helpful post series, but your links to the follow on articles are broken, and there is a 'TODO' in the middle of part 1.