9 ms·
"Let it crash" is largely about restarting systems to known good state in an expanding scope. The Zen of Erlang[1] covers it really well, it's a longer read but
by vvanders 5y ago
"Let it crash" is largely about restarting systems to known good state in an expanding scope. The Zen of Erlang[1] covers it really well, it's a longer read but well worth it if you want to understand a lot of the design choices of Erlang.
Nearly all of the devices you use employ these approaches in one form or another. A watchdog timer[2] is a pretty simple and powerful version of this. Timeouts and retries follow a somewhat similar approach, Erlang just embraces that across the whole language. It really is a fascinating approach to a different design space(latency and reliability over throughput) through the requirements a telecom stack necessitated.
[1] https://ferd.ca/the-zen-of-erlang.html https://ferd.ca/the-zen-of-erlang.html
[2] https://en.wikipedia.org/wiki/Watchdog_timer https://en.wikipedia.org/wiki/Watchdog_timer
- 1propionyl 5y agoPerhaps another succinct phrasing that might be more easily grokked without spending time working on Erlang systems is: "Let it rollback and retry" The way one thinks about processes in Erlang is different than how one thinks about threads in most languages. The expected behavior when you kill a process is that it will be right back up with a known good state very quickly, and you won't have to do much about it yourself because it's handled in a supervisor far from your local process. It's subtle, but it makes a huge difference. In most languages one expects a thrown unhandled exception to wreak havoc. But in Erlang graceful failure and restarting is the norm, not the exception. It's the expected behavior. Moreover, the responsibility of maintaining the process tree integrity is delegated fully to specific processes. "Business intelligence" (to ape a phrase) nodes are very effectively isolated from having to care. If they don't know how to handle such a restart, you just let them crash/be killed and restarted too.
- nickjj 5y ago> "Let it crash" is largely about restarting systems to known good state in an expanding scope. This phrase has always thrown me for a loop in the context of most web development. Mainly because if an application were coded in a way where it's crashing chances are it's never going to get itself back into a working state. For example if your web app throws a 500 because your code is syntactically invalid or is doing something wildly wrong it doesn't matter how many times you restart the web server or spawn another process, it's not going to work. It's going to fail until someone updates the code base to fix the human mistake. Most modern web frameworks can also handle the case where the /oops URL throws a 500 but the home page and everything else works. One page throwing an exception doesn't bring down everything. Now if you're talking about things like retrying a database connection at startup until either a timeout hits or the DB becomes available, that type of stuff is very useful but this is something I've seen included in a lot of web frameworks in a lot of languages. It's essentially a few line while loop that looks for a specific type of exception and then calls the connect function until it works or times out. In general I find in Elixir you're also dealing with error handling on a per function basis because it's common practice to do the ok / error tuple pattern. This is defensive programming to ensure you have an understanding of the system you're developing, just like you would do a try / except in other languages. For example if you were doing token based authentication you'd want your function to return ok and the data you want when it successfully verifies the token but you'd also want to handle the 2 failing cases individually, one error / message for when the token expired and another error / message for when the token was tampered with.
- sodapopcan 5y ago> Mainly because if an application were coded in a way where it's crashing chances are it's never going to get itself back into a working state. "Let it crash" is not about syntax errors, it's about unexpected (ie, exceptional) errors, often as the result of a user taking a completely unexpected path in a large system that was never thought of by programmers (it's probably about more than that but I'm a BEAM n00b). It's happened plenty in web dev for me where a production worker crashes and it simply requires a restart or to be reset back to a known state because the user did something unexpected. As for tuple return in Elixir, that is simply doing it wrong if it is being used for defensive programming (and the antithesis of "let it crash"). It's meant for handling known errors and makes for a concise way of dealing with with it in the functional world—it's similar to Go's multiple returns. e.g. compared to OO (they are both pretty clean by me) # rails foo.update(params) if foo.save do_something else handle_error(foo) end # elixir case Repo.update(foo, foo_args) do {:ok, updated_foo} -> do_something(updated_foo) {:error, error} -> handle_error(error) end It's otherwise very common for an elixir function to return a bare value if it's expected to always work (and "let it crash" if it doesn't).