17 ms·
Crash-Only Software and Recursive Microreboots
- strofcon 5y agoSo... Erlang / OTP? Sweet.
- girvo 5y agoAlternatively, https://mirage.io/ https://mirage.io/ -- sort of, anyway
- heisenzombie 5y agoWell, a bit more radical than that. In Erlang world, I think this is basically advocating that your GenServer "handle_call/3" "handle_cast/2" methods should only ever return "{stop,...}" and never "{reply...}" or "{ok...}"
- macintux 5y agoI don’t read it that way; I think they’re advocating architecting a collection of components and rebooting them at the first sign of trouble. So, basically Erlang. > This cheap form of recovery engenders a new approach to high availability: microreboots can be employed at the slightest hint of failure, prior to node failover in multi-node clusters, even when mistakes in failure detection are likely; failure and recovery can be masked from end users through transparent call-level retries; and systems can be rejuvenated by parts, without ever being shut down.
- macintux 5y agoI found a functional link to one of the papers, and I figured there’d be some mention of Erlang, but nada. No mentions of actors, either, although that’s less directly relevant. https://www.usenix.org/legacy/event/osdi04/tech/full_papers/candea/candea.pdf https://www.usenix.org/legacy/event/osdi04/tech/full_papers/...
- ccvannorman 5y agoIt would be amazing if my OS had this capability. It is frustrating when things get buggy, slow, and crash-y, and it eats up a lot of my attention to wait for the application / system to reboot. If it could be constantly rebooting, even if it was slower overall, the mitigation of sudden bursts of slow would be VERY worth it. Why are we not funding this!?
- colanderman 5y agoMost (all?) of the paper links appear to be dead. Here's some slides on the topic: http://roc.cs.berkeley.edu/retreats/winter_03/slides/candea_crashonly.pdf http://roc.cs.berkeley.edu/retreats/winter_03/slides/candea_... Interesting to see a name put to this. I used this technique (not having heard of it before) when developing a short but finicky and critical cloud-coordinated cross-datacenter disaster recovery consensus algorithm a few years ago. There was a point at which the algorithm recorded its state and a timestamp to local storage, so progress and correctness were guaranteed in the face of a crash. I came to the same conclusions as these researchers -- this state transition did not need to be fast, but did need to be absolutely robust and well tested. So rather than code and test both the "happy path" and the crash-recovery path -- I just called `_Exit(1)` and only coded and tested the crash-recovery path. (It wasn't a perfect confluence of testing space -- my code exited in a way that did not trigger a core dump; whereas most failure modes did. But generally, core dumps bogging down an already-ill system were an issue we had to contend with.) Would be interesting to see a DSL designed around this technique -- a semi-persisted program state of sorts -- with the happy-path automatically generated as an optimization.
- elliottkember 5y agoI once had a router that would get slower and slower. Rebooting it brought it back to life. I had a timer power strip, and set it to power off at 2am for 15 minutes every night. It worked beautifully.
- bitwize 5y agoSometimes "Did ye try tairnin' it off and on again?" really is the best solution.
- macintux 5y agoHere’s a link to the first paper listed: https://www.usenix.org/legacy/event/osdi04/tech/full_papers/candea/candea.pdf https://www.usenix.org/legacy/event/osdi04/tech/full_papers/...
- macintux 5y agoMore functional links are available at http://roc.cs.berkeley.edu/ http://roc.cs.berkeley.edu/ @dang et al, that link would be preferable to the web archive; more of its paper links are still alive.